AI avatars turn a script and a face (plus a voice) into video of someone talking — either rendered once ahead of time, or generated live in a real conversation. Pre-recorded avatars (HeyGen, Synthesia, D-ID) are the mature, polished end of the market: upload a photo, type a script, get a video in minutes. Realtime interactive avatars (Tavus, and HeyGen/D-ID live products) go further, holding an actual back-and-forth conversation on camera with an LLM behind the face. On the open-source side, SadTalker and LatentSync/Wav2Lip cover photo-animation and lip-sync dubbing for free, if you have a GPU. Be honest about the gap: there is no production-ready open-source option yet for realtime interactive avatars — if you need live conversation today, you are choosing between paid proprietary tools. → Quick pick: polished pre-recorded talking-head → HeyGen or Synthesia; a photo into a talking head → D-ID; live interactive conversation → Tavus; free with a GPU → SadTalker (pre-recorded) or LatentSync (lip-sync).
1.1After this chapter you can
→Generate a pre-recorded talking-head video from a photo and a script
→Understand the difference between pre-recorded and realtime interactive avatars
→Know the open-source options (SadTalker, LatentSync/Wav2Lip) and their GPU/setup tradeoffs
→Recognize the consent and deepfake-disclosure issues around using someone's likeness
1.2Which tool makes pre‑recorded videos?
HeyGen and Synthesia are the polished pre‑recorded avatar services – upload a photo, type your script, and receive a finished talking‑head video in minutes.
1.3How do I animate a photo without studio?
SadTalker lets you animate a static photo into a talking head, while LatentSync with Wav2Lip provides lip‑sync dubbing – both run locally on a GPU and are free to use.
2Matrix 6 rows · 6 tools
heygen
synthesia
d-id
tavus
sadtalker
latentsync
Pre-recorded video
yes
yes
yes
no
yes
no
Realtime interactive
yes
no
yes
yes
no
no
Open source / self-host
no
no
no
no
yes
yes
Runs locally
no
no
no
no
yes
yes
Cost
$29-149/mo
$29/mo+
~$4.70/mo+
free 25min · $59+
$0 · GPU
$0 · GPU
Voice cloning / dubbing
yes
yes
no
no
no
yes
3Lessons 7
3.1Create a photo‑based AI avatar in HeyGen
HeyGen’s Photo Avatar feature that turns a single image into an unlimited talking video.
You will have generated a short video of a realistic avatar speaking your script from just one photo.
Open the HeyGen web app and click “Pick an Avatar”.
Select “Photo Avatar” and upload a clear front‑facing photo of yourself.
Enter a short script (up to 200 characters) in the text box and choose any language you prefer.
Press “Generate video” and wait for the platform to render the avatar with lip‑sync.
When the video appears, click download to save the MP4 file locally.
You'll see A downloadable video where the face from your photo moves its lips in time with the spoken script.
Takeaway A single image can be transformed into a lifelike speaking avatar without any camera or editing software.
3.2Make a talking avatar from a single photo
A browser‑based feature that creates a speaking video from a single portrait by analysing facial geometry and syncing lip movements to supplied audio or text.
Generate a short avatar video that lip‑syncs to your script using only an uploaded photo and optional voice input
Select Photo‑to‑Video in the AI avatar generator
Upload a clear portrait image as the source avatar
Provide either a written script or an audio file for the spoken content
Press Generate to create the lip‑synced video
You'll see A video file where the uploaded photo is animated with realistic facial movements that match the supplied speech
Takeaway A single high‑resolution portrait can be turned into a speaking avatar without any filming or manual animation
Check Which two inputs does the Photo‑to‑Video feature require, and what output does it produce when both are provided?
3.3Apply expressive motion to an avatar video
HeyGen’s Motion Engine, which applies preset gesture and facial dynamics to an existing avatar video.
Add coordinated body language that matches a chosen emotional tone to an existing avatar video
Open the project containing your avatar video
Click the Motion Engine control in the right‑hand panel
Choose a preset such as Excited, Welcoming or Calm
Preview and re‑render the video to apply the motion style
You'll see The avatar now moves its head, shoulders and facial expressions in sync with the speech, reflecting the selected energy style
Takeaway Preset motion styles let you quickly add natural‑looking body language without manual animation
Check Which parts of the avatar does a Motion Engine preset move, and what are they synchronised to?
3.4Add expressive gestures and hand movements to the photo avatar
HeyGen’s Avatar IV motion controls that let you customize facial expression, gesture, and upper‑body animation.
You will enhance the previously created avatar video with natural gestures and expressive face dynamics.
Open the video you generated in HeyGen’s AI Studio editor.
In the right‑hand panel, locate “Gesture Control” and select a preset (e.g., “Enthusiastic” or “Calm”).
Toggle “Facial Expression” and choose an emotion that matches your script tone.
Adjust the timing sliders to sync gestures with key sentences in the script.
Re‑render the video by clicking “Generate” again, then download the updated file.
You'll see The avatar now moves its hands and changes facial expressions in rhythm with the spoken words.
Takeaway Fine‑grained motion controls let a static avatar feel like a real presenter delivering a message.
3.5Localise an avatar video
An AI‑driven translation tool that replaces an avatar video's audio track with multilingual speech while keeping the original visuals unchanged.
Create a version of the avatar video that speaks a different language while preserving all visual elements
Click the Translate button in the left sidebar
Select the source video and choose the target language from the list
Edit the auto‑generated script if needed, then confirm to generate the new audio track
You'll see A duplicate video where the avatar’s lip movements stay the same but the spoken language has changed to the selected target language
Takeaway Separating visual content from audio enables efficient multilingual scaling of AI‑generated videos
Check What stays identical between the source video and its translated version, and what changes?
3.6Localise the avatar video into another language
HeyGen’s Video Translation feature that automatically translates speech, subtitles, and lip‑sync to any of 175+ languages.
You will produce a version of your expressive avatar video in a different language with accurate lip‑sync.
In AI Studio, open the expressive video you just saved.
Click “Translate” and pick a target language from the dropdown (e.g., Spanish).
Choose whether to keep the original voice or enable “Voice cloning” for the new language.
Press “Generate translation”; HeyGen will replace the audio, subtitles, and lip movements.
Download the translated video once processing finishes.
You'll see A new video where the avatar speaks the script in the selected language with perfectly synced mouth movements and on‑screen captions.
Takeaway One click can turn a single avatar video into multilingual versions, simplifying global content distribution.
3.7Publish and share your multi‑language avatar videos
HeyGen’s export and sharing workflow that lets you download high‑resolution files or copy a shareable link.
You will make both the original and translated videos publicly accessible for review or distribution.
From the project dashboard, select each video (original and translated).
Click “Download” and choose 1080p resolution to save high‑quality MP4 files.
Alternatively, click “Copy link” to generate a shareable URL for each video.
Paste the links into a short email or document to distribute to teammates.
Verify that the videos play correctly in a web browser.
You'll see Download prompts and functional URLs that open the avatar videos without requiring HeyGen login.
Takeaway HeyGen provides simple export options, enabling rapid distribution of AI‑generated, localized video content.
4You’ll know it worked 43 checkable outcomes in this chapter
✓The avatar appears in the AI studio with accurate lip‑sync and expressive facial animation matching your recording
✓The video shows the avatar displaying the chosen facial expressions synced with the narration
✓You can correctly identify each term in the interface and explain its purpose
✓The cloned avatar appears in the avatar list labeled with your chosen name and plays back your recorded script correctly
✓Play the exported video and see your face moving in sync with the spoken text
✓The avatar appears in the AI Studio preview matching your likeness and voice
✓The avatar appears in the preview pane with your chosen style and can be swapped without errors
✓During editing you can see each avatar’s different viewpoint when swapping them in scenes
43 outcomes in all — one per recipe below.
5FAQ, Tips & How-to 51
one problem, one solution, one action
?FAQEveryone
How do I create a personalized AI avatar that looks and sounds like me?
Record a short 15‑second clip with clear lighting and natural speech, then go to Avatar → New Avatar → Clone a real person (or Digital Clone). Upload the video, optionally add a voice sample, and let HeyGen process it into a digital twin that mimics your appearance, facial movements and voice. Once created you can save the avatar for any future project.
AI-generated
?FAQEveryone
What’s the easiest way to add expressive gestures to my avatar?
In AI Studio click the Motion Engine icon and choose a preset energy style such as excited, welcoming or calm. The engine automatically animates gestures and facial dynamics to match the tone, so you get coordinated movement without manual keyframing. Preview the result and regenerate the video to apply the changes.
AI-generated
?FAQEveryone
Can I translate an existing HeyGen video into another language without re‑recording?
Yes. From the home page click Translate, select Video Dubbing, upload your original MP4 and choose the target language. HeyGen translates the script, generates a new voice in that language and syncs it with the existing visuals, producing a dubbed version you can preview and download.
AI-generated
?FAQEveryone
How do I quickly generate a short promotional clip featuring my avatar?
Use Avatar Shots: go to Avatar → Avatar Shots, select your digital clone, and write a concise prompt describing the scene, setting and outfit changes. Choose portrait or landscape orientation and click Generate; HeyGen synthesizes rapid camera moves and wardrobe swaps into a 15‑second cinematic promo.
AI-generated
?FAQEveryone
How can I create a digital twin of myself to use as an AI presenter?
Record a short 15‑second clip with good lighting and natural speech, then upload it in HeyGen’s Avatar 5 or Digital Clone creator. The platform extracts your facial movements, voice timbre and speaking style to generate a personalized avatar that mimics your look and voice.
AI-generated
?FAQEveryone
What is the easiest way to add expressive gestures to my avatar?
Use the Motion Engine inside AI Studio: click its icon, choose a preset energy style such as excited, welcoming or calm, preview the movement and regenerate the video. The engine automatically animates facial dynamics and gestures without manual keyframing.
AI-generated
?FAQEveryone
Can I translate an existing HeyGen video into another language without re‑recording?
Yes. From the home page click Translate, select Video Dubbing, upload your original MP4, choose the target language and start the process. HeyGen will generate a new voice in that language and sync it with the existing visuals.
AI-generated
?FAQEveryone
How do I produce a full video just by writing a script?
Open Video Agent, paste or type your script into the prompt box, then select the avatar and voice you have saved. Choose a visual style and any brand assets, click Generate (or Auto‑Generate) and HeyGen will auto‑storyboard scenes, apply backgrounds and render the final video.
AI-generated
▸How-toEveryone
Only a 15‑second clip and need a digital double
HeyGen can clone a real person by recording just 15 seconds of video. The system extracts facial movements, gestures and expressions to build an avatar that mimics you across angles and scenes.
Design with AI lets you define a base portrait and then either pick from pre‑made scenes or write prompts to create new outfits, backgrounds and angles. The avatar’s face stays consistent while the surrounding visuals change.
I have a text script and want an avatar to speak it
AI Studio combines a text script (typed or generated via ChatGPT) with an avatar, motion engine, and chosen look to output a finished video in minutes. Advanced settings let you control motion style for more natural movement.
Video Agent automates script writing, scene layout and avatar rendering. By describing topic, audience and tone in one prompt, the system creates a full video with multiple scenes without manual editing.
Synthesia lets you upload photos and voice recordings to generate an avatar that mirrors your appearance and speech. The platform trains the model on your data, producing a lifelike representation you can use in videos.
Synthesia offers a library of over 200 avatars with selectable expressive styles, allowing you to match tone and body language to your script for more engaging content.
Synthesia lets you generate an avatar that looks and sounds like you by uploading a single photo or a short video. The platform processes the media to build a 3D model with voice synthesis, enabling you to produce videos without ever filming again.
Within the editor you can swap an avatar’s clothing and location by describing what you want. Synthesia’s generative AI creates a preview in real time, letting you match branding or context without external design tools.
After finalizing a video, Synthesia can automatically translate both spoken audio and on‑screen text into over 140 languages, producing localized versions in one click. This streamlines global distribution without manual dubbing or subtitle creation.
By selecting the avatar and using the Media tab, you can describe any desired action (e.g., ‘raise hand’, ‘point to chart’). Synthesia’s AI interprets the prompt and animates the avatar accordingly, adding dynamic movement without manual keyframing.
The video defines six essential words (Avatar, Voice, Project, Scene, AI Studio, Credits) that appear throughout HeyGen. Knowing these terms lets you recognize UI elements and plan your videos without getting lost when labels change.
The left‑hand sidebar contains five main sections (Home, Avatar, Brand, Apps, Projects). Understanding this layout lets you jump directly to the tool you need, regardless of UI updates.
Need a fast prototype video from just a text prompt
Quick Create is a shortcut that generates a video from a simple text prompt, bypassing the full editor. It’s useful for rapid prototyping or testing ideas before committing to detailed editing.
Want every video to match my brand colors and logo
The Brand section stores colors, logos, fonts, and templates. By configuring a brand kit once, all future videos can automatically use the same styling, saving time and ensuring consistency.
D-ID Studio lets you build a video featuring a custom AI avatar by selecting or creating an avatar, adding a script, choosing or cloning a voice, and layering visual elements. The platform integrates avatar generation, text‑to‑speech, and scene design in one workflow, making high‑quality avatar videos fast and accessible.
HeyGen offers a searchable library of pre‑made avatars filtered by age, gender and ethnicity. Selecting the right one saves time and ensures your video matches its target audience.
By recording a short 15‑second clip with clear lighting and natural speech, HeyGen learns your facial movements, voice timbre and speaking style, producing a personalized AI clone.
Using the “Design with AI” option lets you describe a character’s appearance, age, gender, ethnicity and style; HeyGen generates multiple visual variations that can be saved and paired with any voice library.
The AI Studio interface separates script, timeline, avatar/media/captions controls, allowing precise tweaks to wording, caption styling and background images generated on‑the‑fly with AI image models.
Want expressive avatar gestures without keyframing
The Motion Engine lets you pick a preset energy style (excited, welcoming, calm) that automatically animates gestures and facial dynamics, giving the avatar more life without manual keyframing.
HeyGen can automatically translate the script, generate a new voice in the target language and sync it with existing visuals, letting you reach multilingual audiences without re‑recording.
The tool lets you upload any portrait and animates it to speak using AI-generated facial movements. It works by mapping the static image onto a 3D model that syncs lip movements with audio, producing realistic talking heads.
D‑ID provides a library of synthetic voices that can be matched to any avatar. By selecting a voice and adjusting style, you control tone, accent, and pacing without recording audio yourself.
Beyond pre‑made presenters, D‑ID can create new avatar images on the fly using a text prompt that feeds a Stable Diffusion model. This lets you design custom characters without external image tools.
Need a video character in a specific outfit and background
Synthesia lets you change an avatar's look by describing the desired costume or setting. The platform generates a preview instantly, allowing rapid iteration without any design skills.
Want a personal presenter that looks and sounds like you
By recording a short video of yourself reading a script, Synthesia extracts facial movements and voice characteristics to build an avatar that looks and sounds like you, enabling fully automated personal presentations.
You record 15 seconds of video and optionally a voice sample, then HeyGen generates a digital twin that mimics your appearance, speech, and motion. The short clip captures enough facial and body cues for Avatar 5 to produce authentic movements.
Video Agent lets you describe the entire video (topic, style, length) in natural language; HeyGen assembles scenes, adds your avatar, and renders the final clip automatically. This bypasses manual editing for quick content production.
Avatar Shots uses Seedance 2.0 to synthesize rapid scene changes, wardrobe swaps, and camera moves based on a single prompt, producing high‑impact short clips without filming.
The translation feature re‑uses your avatar’s visual content while swapping the audio track with AI‑generated speech in the target language, preserving brand consistency across markets.
Synthesia lets you modify an avatar’s outfit, background, and even upload a custom likeness that mimics your voice and appearance. By describing the desired look in plain text, the platform generates a preview instantly, allowing rapid iteration without any video editing skills.
Beyond static speaking, Synthesia can prompt an avatar to execute actions like walking, cooking, or demonstrating objects by adding a simple textual cue. The system maps the cue to motion libraries, creating dynamic, engaging video segments without manual animation.
HeyGen’s Auto Avatar feature lets you build an avatar either from a public library or by uploading reference images to generate a digital twin. The system analyses facial features and maps them onto a 3D model, enabling realistic lip‑sync and expressions.
AutoVoice can clone an existing voice by analyzing a short audio sample, then synthesize speech that matches tone and cadence. This enables the avatar to speak with a natural‑sounding, personalized voice.
Need a video of a custom avatar speaking your script
By entering a text prompt or full script, HeyGen’s Video Agent auto‑storyboards scenes, matches them with the selected avatar and voice, applies visual styles, and renders the final video in minutes.
Need a video presenter that looks and sounds like you
You record a short video of your face and voice, then Synthesia processes it into a personal avatar that mimics both appearance and speech. This works because the platform uses facial mapping and voice cloning to generate a realistic digital presenter.
You can adjust shirt color, add a logo, and select background scenes for ready‑made avatars, ensuring visual consistency with your corporate identity. The platform applies these style changes instantly across the avatar’s appearance.
Recording three personal avatars from different camera angles lets Synthesia switch viewpoints within a single video, creating a more engaging and professional look. Each angle is treated as a separate avatar that can be swapped scene‑by‑scene.
After generating a video outline, you can replace the default presenter with your personal avatar, adjust its size/position, and change background settings. This lets you personalize every scene without re‑rendering the whole script.
Record a short, well‑lit video of yourself speaking naturally for about 15 seconds. HeyGen extracts your voice, facial expressions and gestures to build an avatar that can be reused in any scene.
Video Agent first builds a scene‑by‑scene blueprint, letting you choose avatar, voice, background and style before any rendering occurs, so you can edit each element without re‑generating the whole video.
Long video or podcast needs bite‑size social clips
Upload a full‑length video or podcast; HeyGen analyzes speech and visual cues to identify key statements and emotional peaks, then outputs trimmed, formatted snippets ready for social platforms.
The system infers facial geometry and natural movement patterns from one image, then animates it to deliver any script, useful for mascots or quick spokespersons when full video isn’t available.
Shows how to generate a speaking avatar from a photo, build a realistic digital twin from a short clip, and convert a script into an editable video using HeyGen tools
Shows how to start using HeyGen by launching a quick video, explaining key terminology, and navigating the sidebar
7FAQ 15
How do I turn a personal photo into a talking video with D‑ID?
Upload a portrait (preferably neutral) in Creative Reality Studio, add your script text, choose language, voice and style, then click Generate. The platform maps the image onto a 3D model that syncs lip movements to the audio and produces an MP4 you can download.
Can I use D‑ID’s built‑in voices instead of recording my own?
Yes. After uploading your image, go to the Voice section, browse the synthetic voice library, preview each option, pick one, optionally adjust its style (e.g., excited or friendly), enter your script and generate the video.
What if I don’t have a suitable photo—can D‑ID create an avatar from text?
D‑ID can generate a new avatar image using a text prompt that runs through a Stable Diffusion model. In the studio, click “Generate AI presenters,” type a descriptive prompt, select resolution, and submit; the result is a fresh avatar ready for script and voice addition.
How do I create a personalized AI avatar that looks and sounds like me?
Record a short 15‑second clip with clear lighting and natural speech, then go to Avatar → New Avatar → Clone a real person (or Digital Clone). Upload the video, optionally add a voice sample, and let HeyGen process it into a digital twin that mimics your appearance, facial movements and voice. Once created you can save the avatar for any future project.
What’s the easiest way to add expressive gestures to my avatar?
In AI Studio click the Motion Engine icon and choose a preset energy style such as excited, welcoming or calm. The engine automatically animates gestures and facial dynamics to match the tone, so you get coordinated movement without manual keyframing. Preview the result and regenerate the video to apply the changes.
Can I translate an existing HeyGen video into another language without re‑recording?
Yes. From the home page click Translate, select Video Dubbing, upload your original MP4 and choose the target language. HeyGen translates the script, generates a new voice in that language and syncs it with the existing visuals, producing a dubbed version you can preview and download.
How do I quickly generate a short promotional clip featuring my avatar?
Use Avatar Shots: go to Avatar → Avatar Shots, select your digital clone, and write a concise prompt describing the scene, setting and outfit changes. Choose portrait or landscape orientation and click Generate; HeyGen synthesizes rapid camera moves and wardrobe swaps into a 15‑second cinematic promo.
How can I create a digital twin of myself to use as an AI presenter?
Record a short 15‑second clip with good lighting and natural speech, then upload it in HeyGen’s Avatar 5 or Digital Clone creator. The platform extracts your facial movements, voice timbre and speaking style to generate a personalized avatar that mimics your look and voice.
What is the easiest way to add expressive gestures to my avatar?
Use the Motion Engine inside AI Studio: click its icon, choose a preset energy style such as excited, welcoming or calm, preview the movement and regenerate the video. The engine automatically animates facial dynamics and gestures without manual keyframing.
Can I translate an existing HeyGen video into another language without re‑recording?
Yes. From the home page click Translate, select Video Dubbing, upload your original MP4, choose the target language and start the process. HeyGen will generate a new voice in that language and sync it with the existing visuals.
How do I produce a full video just by writing a script?
Open Video Agent, paste or type your script into the prompt box, then select the avatar and voice you have saved. Choose a visual style and any brand assets, click Generate (or Auto‑Generate) and HeyGen will auto‑storyboard scenes, apply backgrounds and render the final video.
How do I create a personalized AI avatar that looks and sounds like me?
Record a short, well‑lit 15‑second video of yourself speaking naturally and upload it via the Avatar 5 or Digital Clone creator. The system extracts your facial movements, voice timbre and speech style to build a digital twin. After processing you can review the generated avatar, select the best version and save it for any future project.
Can I add expressive gestures to my avatar without manually animating each movement?
Yes, use the Motion Engine in AI Studio. Choose a preset energy style such as excited, welcoming, or calm, which automatically applies coordinated facial dynamics and hand gestures. Preview the motion, adjust if needed, then regenerate the video to apply the changes.
How can I translate an existing avatar video into another language?
From the home page click Translate, select Video Dubbing, and upload your original MP4. Choose the target language; HeyGen will translate the script, generate a new voice in that language, and sync it with the existing visuals. When processing finishes you can preview and download the dubbed version.
What’s the quickest way to produce a short video using an avatar?
Use Quick Create from the Home screen: enter a brief text prompt or script, pick a stock avatar if needed, and submit. HeyGen generates a one‑minute prototype video automatically, which then appears under Projects for you to review or edit further.
8Glossary 24 terms
Show the 24 terms
HeyGen
Avatar
The on‑screen digital presenter that speaks and moves in your video.
AI Studio
The main editing workspace where you adjust script, captions, background and avatar settings.
Motion Engine
A feature that adds preset expressive gestures and facial movements to an avatar automatically.
Video Agent
A tool that builds a video from a text description by selecting scenes, avatars, voices and styles before rendering.
Avatar Shots
A function that creates short cinematic clips with rapid scene changes and outfit swaps from a single prompt.
Seedance 2.0
The underlying technology used by Avatar Shots to synthesize dynamic camera moves and wardrobe changes.
Nano Banana
An AI image model you can call on to generate custom background images for your video.
AutoAvatar
A feature that generates a personalized avatar from uploaded photos or a short video of yourself.
AutoVoice
A function that clones a voice by analyzing a brief audio sample and creates a synthetic speech model.
Credits
The internal currency you spend to generate avatars, voices or videos within HeyGen.
Brand Kit
A collection of your logo, colors, fonts and templates that can be applied automatically to new videos for consistent branding.
Quick Create
A shortcut that produces a simple video from a brief text prompt without opening the full editor.
Synthesia
Personal Avatar
A lifelike AI copy of you that can appear and speak in Synthesia videos.
Avatars → Customize
The menu path to edit an avatar’s colors, logo, clothing or background.
hex code
A six‑digit combination (like #FF5733) that specifies an exact color for branding.
Enterprise plan
The paid subscription level that lets you upload a custom logo onto avatars.
Avatar dropdown
A list in the video editor where you pick which avatar (or angle) appears in a scene.
Change All
An option that replaces the default presenter with your chosen avatar across every scene at once.
Add Space
A button that lets you select or upload a background image for an avatar’s scene.
Custom Avatar
An avatar whose outfit, setting and voice are defined by typing a description in plain text.
Action Prompt
A field where you type a short command (e.g., “walk across the screen”) to make the avatar perform that motion.
Media tab
The panel on the right side of the editor used for entering natural‑language action descriptions.
Translate (in export)
A one‑click option that creates separate video files with dubbed audio and subtitles in chosen languages.
Multi‑Angle Avatars
Three separate personal avatars recorded from different camera positions that can be swapped to change viewpoint within a single video.