Twin model / avh-01
Workflow / 7 steps
Ready / awaiting capture
Guide / 7 steps / checked Sep 2026
Create an AI clone of yourself that records your videos for you
All it needs is a 2-minute phone clip. Below: the exact HeyGen workflow, what each step costs in credits, and where the clone still gives itself away.
Affiliate link: we earn 20% of what you pay HeyGen in your first 12 months. The free plan costs nothing and includes one custom avatar.
Step 01 / Film the training clip
Record two minutes of you, talking
One uninterrupted take. HeyGen's filming guide asks for:
- At least 2 minutes of continuous speech. Avatar V can start from 15 seconds, but more footage gives it more of your motion to learn.
- 1080p minimum. 4K at 60 fps is better.
- A quiet room, good even light, and clothes without busy patterns or large logos.
- Eye contact with the lens, head turns under 30 degrees, hands below chest level.
- A 2 to 3 second pause at the start, and 1 to 2 second pauses between longer sentences, lips closed.
- On a phone, press and hold your face on screen until the lock icon appears. That locks focus and exposure.
Step 02 / Record the consent clip
Prove the face is yours
Before it trains anything, HeyGen asks for a short consent video, and it must show the same person as the training footage. It does not accept someone else's video without permission, or footage generated by other AI tools.
That rule is the safeguard against anyone cloning a person from public videos. It also means a team member's clone needs that team member in front of the camera.
Same person / on camera / every time
Step 03 / Train with Avatar V
Pick the model that learns your motion
Choose Avatar V, HeyGen's newest model. HeyGen says it conditions on your whole video rather than a single frame, which is how it keeps your gestures and expressions consistent across scenes.
Two catches: it is HeyGen's most expensive model to render, and each Avatar V video is capped at 3 minutes.
48 credits per minute / 3 min per video
Step 04 / Clone your voice
A real face with a stock voice looks fake
The mismatch is the first thing viewers notice. The Creator plan includes one voice clone and Pro expands voice usage. Give it clean audio: a quiet room and a microphone close to your mouth matter more than the camera does.
Creator: 1 voice clone
Step 05 / Write for the ear
The script is now the whole production
Your twin reads exactly what you type, so the writing does the work a camera crew used to.
- About 150 spoken words is roughly one minute.
- Short sentences, one idea each. Read it aloud once before you paste it.
- Spell numbers, brands and acronyms the way you say them.
- Keep Avatar V pieces under 3 minutes and split longer scripts into chapters.
Step 06 / Render and check
Watch one take at full size, with sound
Credits are spent when you render, so check one before you batch ten. Look for:
- Lip-sync on fast consonants (p, b, m) and on names.
- Hands and gestures near the edge of the frame.
- The eye line: does it hold the lens the way you did?
- Brand names. Fix mispronunciations with phonetic spelling in the script.
Step 07 / Translate
Multiply one good video into 175+ languages
Translation is the cheapest multiplier HeyGen has: 6 credits a minute with lip-sync, 10 in precision mode, 4 without lip-sync. That is an eighth of the cost of rendering new Avatar V footage.
Have a native speaker watch the first one. Some reviewers say translations are not always smooth.
6 credits per minute / 175+ languages