Heidelberg AICurriculum
Track 5 · Intermediate

Create Media

voice, avatars & video

This track teaches you how to create media that people can watch and hear. It is for anyone who wants to turn text, images, or audio into voiceovers, avatars, pictures, or video without relying on proprietary services. When you finish the chapters you will be able to give a script a face using AI avatars, generate pictures from prompts with both cutting‑edge and free local tools, create short video clips from sentences or photos, add spoken dialogue to an app and make it understand speech, and build real‑time voice agents that converse live. Start with Chapter 1 on AI avatars, then move to Chapter 2 for image generation, followed by Chapter 3 on video creation. After you have visual media covered, continue with Chapter 4 to add text‑to‑speech and speech‑to‑text capabilities, and finish with Chapter 5 for real‑time voice agents. If you are short on time, you can skip the optional deep dive into frontier image apps in Chapter 2 and go straight to the free, local route before proceeding. Follow this order to build a solid foundation before adding interactive voice features, ensuring each skill builds on the previous one.

5 chapters 603 recipes Intermediate 0/5 done
Comparison matrix from AI avatars

AI avatars turn a script and a face (plus a voice) into video of someone talking — either rendered once ahead of time, or generated live in a real conversation. Pre-recorded avatars (HeyGen, Synthesia, D-ID) are the mature, polished end of the market: upload a photo, type a script, get a video in minutes. Realtime interactive avatars (Tavus, and HeyGen/D-ID live products) go further, holding an actual back-and-forth conversation on camera with an LLM behind the face. On the open-source side, SadTalker and LatentSync/Wav2Lip cover photo-animation and lip-sync dubbing for free, if you have a GPU. Be honest about the gap: there is no production-ready open-source option yet for realtime interactive avatars — if you need live conversation today, you are choosing between paid proprietary tools. → Quick pick: polished pre-recorded talking-head → HeyGen or Synthesia; a photo into a talking head → D-ID; live interactive conversation → Tavus; free with a GPU → SadTalker (pre-recorded) or LatentSync (lip-sync).

heygen
synthesia
d-id
tavus
sadtalker
latentsync
Pre-recorded video
yes
yes
yes
no
yes
no
Realtime interactive
yes
no
yes
yes
no
no
Open source / self-host
no
no
no
no
yes
yes
Runs locally
no
no
no
no
yes
yes
Cost
$29-149/mo
$29/mo+
~$4.70/mo+
free 25min · $59+
$0 · GPU
$0 · GPU
Voice cloning / dubbing
yes
yes
no
no
no
yes
The chapters outline order · updated · your progress · recipe load