Making Human Voices

Clone a voice with human prosody

in voice cloning

Meet Soul TTS, leading voice cloning with human-like prosody.

Soul TTS

The reference speaker sets the voice identity. All models generate the target text shown below.

For each reference speaker and target text pair, we select the first result from every model (all models use default recommended settings).

Soul TTS captures reference speaker qualities and generates correct semantics with high fidelity. Per 30 seconds of audio, generation takes ~1.2s, ~14s, and ~50s respectively for Soul TTS, Higgs Audio, and VibeVoice on an A100.

These prompts are stylistically in-distribution for Soul TTS. Better samples for the other models could likely be obtained via prompt or hyperparameter optimization. More optimized implementations may also yield speed-ups.

Compare against leading open-source TTS models:

Sample 1
Reference speaker
Target text

“Dobby has to tell you something, sir. Something terrible. You must not trust this place. Dobby was bound once, bound forever to serve. But no longer. Now, Dobby has to use his freedom to keep you safe.”

VibeVoice (latest)
Soul TTS
Higgs Audio (latest)
Sample 2
Reference speaker
Target text

“Messi picks it up just inside the half, he turns away from one, he glides past another and the whole stadium can feel it coming! He's driving at the heart of the defense now, nobody can get near him. Oh! He's curled it into the top corner and this place has lost its mind! Absolutely magical from the little genius!”

VibeVoice (latest)
Soul TTS
Higgs Audio (latest)
Sample 3
Reference speaker
Target text

“So I'm standing there trying to act normal, right, and I reach for a drink and somehow knock the entire tray out of the waiter's hands and everyone just turns and stares at me. And I'm thinking, you know, please just let the floor open up and swallow me whole, but no of course not, the night was just getting started.”

VibeVoice (latest)
Soul TTS
Higgs Audio (latest)

Architecture

The model learns to denoise latents into speech that matches both the speaker's voice and the text content.

Reference audio conditions the decoder through cross-attention.

Audio Input ≤120s @ 44.1kHz → DAC-VAE encode → continuous latents 44.1kHz → ~86Hz Speaker Encoder Causal Transformer patch_size=4 2560 → 640 seq len Speaker embedding K, V projection Text Encoder Bidirectional Transformer UTF-8 bytes, max 768 Diffusion Decoder Joint Self-Cross Attention K,V = concat(self, speaker, text) QK-norm · Gated attention SwiGLU MLP AdaLN · LoRA adapters Rectified Flow Noise Latent Audio 30 steps DAC-VAE Decode continuous latents → DAC-VAE decoder 44.1kHz output Audio
scroll

With thanks to Jordan Darefsky (Anthropic)

Diffusion process

Rectified flow learns straight paths from noise to data.

t=1.0 (noise)
t=0.75
t=0.5
t=0.25
t=0 (audio)
Sampling 30 steps
Sampler Euler ODE
CFG scale 3.0 (joint)
RTF <0.05

Block-wise streaming

First audio arrives before generation finishes.

Generate and play in parallel.

Generate
Playback
Latent prefix encoder Variable block size
TTFB optimization Shorter initial blocks

Attention

The denoiser sees itself, the speaker, and the text simultaneously.

One operation over all context.

Q (denoiser) K, V sources
Self
Speaker
Text
attn = softmax(Q @ K.T / √d) @ V K,V = cat(self_kv, spk_kv, txt_kv)

Training

160K Hours of audio
800K Steps
768 Batch size
TPU v4-64 Compute
Optimizer Muon
Precision BF16
Dropout 10% (spk/txt independent)
Transcription WhisperD (diarization + events)

Classifier-free guidance

Amplify the signal.

Joint 2× NFE
ε = ε_uncond + s(ε_cond − ε_uncond)
Independent 3× NFE
Separate speaker and text scales
Alternating 2× NFE
Temporal score rescaling

Audio codec

DAC-VAE gives native continuous latents.

Input 44.1kHz PCM
→
DAC-VAE encoder Continuous latents
→
Model latents Diffusion target