Twin model / avh-01
Point cloud / live
Idle / awaiting capture
Independent guides / checked Sep 2026
Your face. Your voice. Without the camera.
AI avatars now learn you from a short phone clip, then present anything you type in 175+ languages. We read the fine print behind the demos, so you know what you get before you pay.
Or start on HeyGen's free plan. We earn 20% of what you pay HeyGen in your first 12 months if you subscribe through our links. It never changes what we publish, including the parts HeyGen would rather we skipped.
Step 01 / Capture
Film two minutes of yourself on your phone
HeyGen's filming guide asks for one uninterrupted two-minute take in 1080p or better, starting with a two to three second pause. Its newest model, Avatar V, can start from just 15 seconds. A phone or a webcam is fine.
Minimum 15 s / recommended 2 min
Step 02 / Render
Type a script. Your twin performs it.
The avatar delivers your script with your face and gestures, and with a voice clone, your voice. Avatar V clips run up to three minutes each, and each minute costs 48 credits. That cost is the number to plan around.
$29 plan / about 12 min of Avatar V
Step 03 / Translate
Publish the same video in 175+ languages
Video translation re-voices a finished video and re-syncs the lips. At 6 credits a minute it costs far less than generating new avatar footage, which makes it the cheapest way to multiply one good video.
6 credits per minute with lip-sync