model

XTTS-v2

XTTS‑v2 is an open‑source text‑to‑speech model that can clone a speaker’s voice from just six seconds of audio and synthesize speech in 17 languages. The model is hosted on Hugging Face and can be run locally via the Coqui TTS Python library, with optional GPU acceleration. It is released under the Coqui Public Model License and has no monetary cost.

  • free
  • self-host
  • Commercial

What it can help with

  • text-to-speech
  • voice cloning
  • multilingual
  • coqui
  • local inference