Create Media
voice, avatars & video
This track teaches you how to create media that people can watch and hear. It is for anyone who wants to turn text, images, or audio into voiceovers, avatars, pictures, or video without relying on proprietary services. When you finish the chapters you will be able to give a script a face using AI avatars, generate pictures from prompts with both cutting‑edge and free local tools, create short video clips from sentences or photos, add spoken dialogue to an app and make it understand speech, and build real‑time voice agents that converse live. Start with Chapter 1 on AI avatars, then move to Chapter 2 for image generation, followed by Chapter 3 on video creation. After you have visual media covered, continue with Chapter 4 to add text‑to‑speech and speech‑to‑text capabilities, and finish with Chapter 5 for real‑time voice agents. If you are short on time, you can skip the optional deep dive into frontier image apps in Chapter 2 and go straight to the free, local route before proceeding. Follow this order to build a solid foundation before adding interactive voice features, ensuring each skill builds on the previous one.
AI avatars turn a script and a face (plus a voice) into video of someone talking — either rendered once ahead of time, or generated live in a real conversation. Pre-recorded avatars (HeyGen, Synthesia, D-ID) are the mature, polished end of the market: upload a photo, type a script, get a video in minutes. Realtime interactive avatars (Tavus, and HeyGen/D-ID live products) go further, holding an actual back-and-forth conversation on camera with an LLM behind the face. On the open-source side, SadTalker and LatentSync/Wav2Lip cover photo-animation and lip-sync dubbing for free, if you have a GPU. Be honest about the gap: there is no production-ready open-source option yet for realtime interactive avatars — if you need live conversation today, you are choosing between paid proprietary tools. → Quick pick: polished pre-recorded talking-head → HeyGen or Synthesia; a photo into a talking head → D-ID; live interactive conversation → Tavus; free with a GPU → SadTalker (pre-recorded) or LatentSync (lip-sync).
- 5.1 AI avatars Give your script a face — pre-recorded or live 2026-08-06 51
- 5.2 AI image generation Turn a prompt into a picture — frontier apps vs. the free, local route 2026-08-06 314
- 5.3 AI video generation Turn a sentence — or a photo — into a video clip 2026-08-06 27
- 5.4 Voice: text ↔ speech Give your app a voice, or teach it to listen 2026-08-06 203
- 5.5 Realtime voice agents Talk to AI out loud — and have it talk back, live 2026-08-06 8