How does voice cloning work?

Soul TTS is Soul’s own text-to-speech model. It clones a voice from reference audio, then generates new speech in that voice.

A reference speaker sets the voice identity. The model generates any target text in that voice, with the rhythm, emphasis, and emotional delivery of real human speech.

Why Soul builds its own TTS

Companions live or die on how they sound. Flat narration breaks the feeling of talking to a person. Soul TTS is built for prosody first.

On an A100 it generates 30 seconds of audio in about 1.2 seconds, compared with roughly 14 seconds for Higgs Audio and 50 seconds for VibeVoice under the same conditions. The architecture write-up, with audio comparisons, is Making Human Voices.

In the app

To clone or change a voice, see How do I choose or clone a voice?.

Was this article helpful?