Voxtral TTS | Summary
26 Mar 2026 | Paper Review Speech Generation Voice Cloning Flow Matching Direct Preference OptimizationContents
- Summary
- 1 Introduction
- 2 Modeling
- 3 Training
- 4 Results
- 5 Analysis
- 6 Inference and Serving in vLLM-Omni
- 7 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Voxtral TTS.
- 2026-03-26 (arXiv)
- Mistral-AI, :, Liu, Alexander H., Tacnet, Alexis, Ehrenberg, Andy, Lo, Andy, Sun, Chen-Yo, Lample, Guillaume, Lagarde, Henry, Delignon, Jean-Malo, et al.
- Paper
- Project Page
Summary
- Voxtral TTS is a multilingual zero-shot text-to-speech model supporting 9 languages and reference audio as short as 3 seconds. Its hybrid architecture generates semantic speech tokens autoregressively and acoustic tokens through a lightweight, frame-conditioned flow-matching transformer.
- Voxtral Codec represents 24 kHz mono speech at 12.5 Hz using one ASR-distilled semantic VQ token and 36 acoustic FSQ tokens per frame, at approximately 2.14 kbps. Training combines semantic cross-entropy and acoustic flow matching, followed by Direct Preference Optimization adapted to both discrete and continuous outputs.
- Blind human evaluations report a 68.4% voice-cloning win rate and a 58.3% implicit-steering flagship-voice win rate against ElevenLabs Flash v2.5. Voxtral TTS achieves higher automatic speaker similarity than both evaluated ElevenLabs models on every reported benchmark row, but does not consistently achieve lower WER and loses the flagship-voice comparisons against Gemini 2.5 Flash TTS.
- With CUDA graphs and asynchronous streaming in vLLM-Omni, a single NVIDIA H200 serves 32 concurrent requests with 552 ms time to first audio, RTF 0.302, and 1430.78 char/s/GPU under the reported 500-character input and 10-second reference conditions. The weights are released under CC BY-NC; these serving results do not establish performance on other hardware or workloads.
The aggregate comparison favors Voxtral TTS more strongly for voice cloning than for flagship voices: 68.4% versus 58.3% against ElevenLabs Flash v2.5. The flagship setup uses 77 text examples, while cloning uses short references and 60 prompts; these are distinct evaluation conditions rather than a controlled test of generalization.
1 Introduction
The research problem is to preserve speaker identity, naturalness, and expressivity in zero-shot TTS without generating every acoustic codebook autoregressively. Building on codecs that separate semantic and acoustic information, Voxtral TTS assigns temporal sequence modeling to an autoregressive backbone and dense acoustic generation to a conditional continuous model.
- ASR-distilled semantic tokens provide a discrete stream for long-range modeling, while FSQ acoustic dimensions support continuous flow matching followed by discretization.
- The paper adapts DPO to this hybrid output space and evaluates intelligibility, predicted naturalness, speaker similarity, human preference, and streaming performance.
2 Modeling
A reference voice is encoded by Voxtral Codec and concatenated with text tokens as input to the decoder backbone. At each generated frame, a linear head predicts a semantic token and a flow-matching transformer predicts acoustic values from the backbone hidden state; the resulting discrete tokens are fed back into the backbone and decoded into a waveform.
- Figure 2 illustrates reference audio of 3s–30s, a 12.5 Hz audio-frame rate, and generation ending at the special <EOA> token.
- Each frame corresponds to 80 ms of audio and contains one semantic token plus 36 acoustic tokens.
The diagram separates temporal generation from acoustic detail: reference audio frames and text enter one autoregressive backbone, whose hidden states support semantic prediction and repeated acoustic flow evaluations. Completed audio frames are fed back into the sequence until <EOA>, illustrating why the architecture remains autoregressive across time but not across acoustic codebooks.
2.1 Voxtral Codec
Waveform Autoencoder. Voxtral Codec is a causal convolutional–transformer autoencoder with approximately 300M parameters. It divides 24 kHz mono audio into non-overlapping patches of 240 samples, projects the resulting 100 Hz sequence to 1024-dimensional embeddings, and compresses it to 12.5 Hz, 292-dimensional latents.
- Each of four encoder blocks contains a 2-layer causal transformer and a causal CNN. The first three CNNs downsample by 2×; the fourth uses stride 1 to project into the latent space.
- Attention windows decrease as 16 → 8 → 4 → 2. The transformers use ALiBi, QK-norm, and LayerScale initialized at 0.01.
- The decoder reverses the compression using causal projections, transposed convolutions, and causal transformers before reconstructing 240-sample waveform patches. Table 1 specifies that the decoder reverses the encoder’s stride sequence.
Representation Quantization. The latent is divided into a 256-dimensional semantic component and a 36-dimensional acoustic component. The former uses an 8192-entry learned VQ codebook; each acoustic dimension is independently quantized into 21 uniform FSQ levels after a tanh activation.
- Semantic VQ is applied to 50% of training samples, with the remainder left unquantized.
- Acoustic training uses 50% FSQ-quantized samples, 25% with uniform noise of magnitude 1/L where L=21, and 25% unquantized samples.
- The reported bitrate is (12.5\times(\log_2 8192+36\times\log_2 21)\approx 2.14\,\mathrm{kbps}).
Semantic Token Learning. A frozen supervised Whisper model supplies decoder hidden states and cross-attention weights for an auxiliary ASR distillation loss. Projected post-VQ semantic embeddings are softly aligned to Whisper’s last-layer decoder states and trained using cosine distance.
- The alignment uses cross-attention heads associated with word timestamps through DTW, followed by normalization across decoder tokens, median filtering, head averaging, and interpolation to the codec’s 12.5 Hz frame rate.
- This procedure does not require an external forced aligner or paired transcripts for codec training, although it relies on a pretrained supervised ASR model.
- The authors motivate continuous hidden-state targets as richer supervision than hard transcript labels, including confidence and phonetic similarities.
Adversarial Training. Eight multi-resolution STFT discriminators distinguish original from reconstructed audio using a hinge loss. The codec is optimized with an L1 feature-matching loss over discriminator activations rather than the standard GAN generator loss.
- The STFT sizes are 2296, 1418, 876, 542, 334, 206, 126, and 76; Table 1 reports 256 discriminator channels.
- The stated motivation is that evolving discriminator features supply an increasingly discriminative reconstruction signal.
Training Objective. The codec objective is αLfeature + βLASR + γLL1 + γLSTFT + δLcommit, with α=1.0, β=1.0, (\gamma=0.9999^t), and δ=0.1, where t is the training step. Waveform and STFT-magnitude reconstruction losses therefore decay together, while feature matching, ASR distillation, and VQ commitment retain fixed weights.
- The commitment term is Lcommit = ||ze − sg(zq)||²₂, where sg denotes stop-gradient.
- Table 1 consolidates the architecture and quantization settings. The paper states that design choices were ablated, but does not provide those codec ablation results.
The semantic and acoustic branches share the waveform encoder and decoder but use different quantizers and supervision. ASR distillation acts specifically on the semantic branch, whereas adversarial feature-matching and reconstruction signals train the waveform reconstructed from both branches.
The configuration combines 240-sample patches and three stride-2 reductions to obtain a 12.5 Hz bottleneck from a 100 Hz patch sequence. Its 8192-entry semantic codebook and 36 × 21 acoustic quantization structure define the discrete interface used by the generator; the footnotes specify stochastic quantization and training-stability settings.
2.2 Decoder Backbone
The decoder-only backbone follows Ministral 3B. Each audio frame is embedded by summing 37 codebook-specific lookup vectors, and each output hidden state is projected to logits over 8192 semantic entries plus <EOA>.
- Semantic prediction uses cross-entropy; the same hidden state conditions acoustic flow matching.
- Flow-generated acoustic values are quantized before the next autoregressive step, maintaining a fully discrete input interface.
2.3 Flow-Matching Transformer
The acoustic generator is a bidirectional 3-layer transformer with the same width as the backbone. At each frame it processes three inputs—backbone hidden state h, sinusoidally embedded integration time t, and a 36-dimensional acoustic state xt—using separate projections.
- Conditioning is dropped with 10% probability during training. Inference uses Euler integration with 8 function evaluations, Δt=1/8, and CFG α=1.2.
- The guided velocity is vt = αvθ(xt,t,h) + (1−α)vθ(xt,t,∅), and the integration update is (x_{t+\Delta t}=x_t+v_t\cdot\Delta t). The unconditional branch uses a zero vector in place of h.
- CFG runs only the small acoustic transformer twice, not the full backbone. Outputs are quantized to 21 FSQ levels before feedback.
- The paper reports better results than AdaLN conditioning and greater expressivity than MaskGIT and Depth Transformer alternatives, without numerical comparisons. The stated per-frame costs are sequence length 38 for MaskGIT and 36 autoregressive steps for Depth Transformer, versus three input positions and 8 NFEs for flow matching; these counts are not a measured runtime comparison.
3 Training
Training has two stages: paired audio–text pretraining for semantic and acoustic generation, followed by preference optimization using generated winner–loser pairs. The post-training stage also retains the pretraining objective on high-quality speech; the authors report that longer training on synthetic preference data produced more robotic speech.
3.1 Pretraining
Pretraining samples have the form (A1, T2, A2), where A1 is a same-speaker reference, T2 is the pseudo-labelled transcript of target audio A2, and the segments are separated by <next> and <repeat>. Loss is applied only to A2, combining semantic cross-entropy with a squared-error conditional flow-matching objective.
- Transcripts are produced by Voxtral Mini Transcribe. A1 and A2 need not be temporally adjacent; both have maximum duration 180 seconds, and A1 must be at least 1 second. The paper reports best behavior for reference prompts between 3 and 25 seconds.
- The backbone is initialized from Ministral 3B, while acoustic flow matching, audio embeddings, and output projections are randomly initialized. Text embeddings are frozen to improve robustness to infrequent transcription tokens.
- VAD-based weighting reduces the influence of non-speech frames and assigns zero weight to extremely long silences. LLM transcript rewrites improve robustness to normalized and unnormalized forms such as “5 - 4” and “five minus four”.
- Equations (6)–(7) define the velocity target as x1 − x0 and label x0 as data and x1 as Gaussian noise; this reverses the endpoint labels used in Section 2.3.
3.2 Direct Preference Optimization
DPO combines a standard semantic-token preference objective with a flow-based acoustic preference objective. The acoustic objective compares summed winner and loser velocity-prediction errors relative to a reference model, with an independently sampled integration time at each sequence position.
- The policy and reference use consistent sampled times and x0 values at corresponding sequence locations. Normalizing by the winner’s length was found to cause instability.
- Semantic and acoustic losses have equal outer weights, with βsemantic=0.1, βacoustic=0.5, and learning rate 8e−8.
- Held-out single-speaker references and texts synthesized by Mistral Small Creative 1 drive rejection sampling. Pair selection uses WER, speaker similarity, loudness consistency, UTMOS-v2, and other LM-judge metrics.
- Post-training lasts 1 epoch alongside the high-quality-speech pretraining objective. Longer training on synthetic data produced more robotic speech in the authors’ observations.
4 Results
The results separate codec reconstruction quality from end-to-end TTS quality. Codec testing uses Expresso, automatic TTS evaluation uses SEED-TTS and the nine supported MiniMax-TTS languages, and blind human evaluations examine flagship-voice emotion steering and zero-shot voice cloning.
4.1 Voxtral Codec
On Expresso, Voxtral Codec outperforms Mimi–16cb on every reported objective quality metric at similar bitrate: Table 2 reports 2.1 versus 2.2 kbps, PESQ 3.05 versus 2.67, ASR-WER 10.66% versus 11.01%, and speaker similarity 0.843 versus 0.829. This supports a reconstruction-quality advantage at the compared bitrate, not superiority over every Mimi configuration.
- Voxtral Codec also has lower Mel and STFT distances than Mimi–full 32cb, but that 4.4 kbps configuration has higher PESQ, ESTOI, and speaker similarity and lower ASR-WER. All compared configurations operate at 12.5 fps.
- Here, ASR-WER compares ASR transcriptions of the source and reconstructed audio rather than measuring generated TTS against the requested text.
- The paper reports an internal subjective assessment of comparable or better speech quality than Mimi–16cb, without providing listener counts or preference statistics.
At similar bitrates, Voxtral Codec improves every listed quality metric over Mimi–16cb, including PESQ from 2.67 to 3.05 and speaker similarity from 0.829 to 0.843. The full 32-codebook Mimi configuration retains advantages in several perceptual and intelligibility metrics at roughly twice the bitrate, so the result is principally a matched-bitrate comparison.
4.2 Automatic Evaluations
Automatic evaluation measures WER with Voxtral Mini Transcribe v2, predicted MOS with UTMOS-v2, and reference-to-generation cosine similarity with ECAPA-TDNN speaker embeddings. Table 3 shows Voxtral TTS leading speaker similarity on all nine MiniMax-TTS languages and SEED-TTS, while WER leadership varies by language and comparator.
- For MiniMax English, Voxtral scores WER 0.63%, UTMOS 4.30, and speaker similarity 0.786; ElevenLabs v3 scores 0.48%, 4.27, and 0.484, and ElevenLabs Flash v2.5 scores 0.33%, 4.27, and 0.489.
- For SEED-TTS, Voxtral scores 1.23%, 4.11, and 0.628; ElevenLabs v3 scores 1.26%, 3.92, and 0.392, and Flash v2.5 scores 0.86%, 4.09, and 0.413.
- The paper observes that Flash v2.5 often scores better automatically than ElevenLabs v3, although v3 fares better in human emotion-steering evaluations. UTMOS is described as weakly correlated with preference and poorly calibrated across languages.
Voxtral leads speaker similarity across all reported tasks, but the WER rankings vary and its UTMOS advantage is not universal. For example, English speaker similarity is 0.786 versus 0.484 and 0.489, while English WER is higher than both ElevenLabs baselines; identity preservation and intelligibility are distinct outcomes.
4.3 Human Evaluations
Human evaluations compare unidentified model outputs and allow judgments of “slightly better”, “much better”, “both good”, or “both bad”. All audio, including references, is resampled to 24 kHz WAV to reduce format-related bias.
- The flagship-voice evaluation uses 77 prompts: 11 neutral and 66 associated with an expected emotion.
- Figure 1 identifies 60 text prompts for the voice-cloning comparison. Its aggregate preference bars show 58.3% for flagship voices and 68.4% for voice cloning against ElevenLabs Flash v2.5.
4.3.1 Flagship voices. Voxtral’s British-Female, British-Male, American-Male, and French-Female voices are compared with competitor voices of matching gender and accent. Three native speakers of the same dialect annotate each pair, and Table 4 reports win rates excluding ties.
- Explicit steering uses free-form emotion instructions for Gemini 2.5 Flash TTS, bracketed emotion tags for ElevenLabs v3, and an emotion-bearing same-speaker reference for Voxtral TTS, which does not support emotion tags or text instructions.
- Under explicit steering, Voxtral wins 51.0% against ElevenLabs v3 and 35.4% against Gemini 2.5 Flash TTS.
- Implicit steering supplies no emotion label; Voxtral uses a neutral reference. Its win rates are 58.3% against Flash v2.5, 55.4% against ElevenLabs v3, and 37.1% against Gemini 2.5 Flash TTS.
Excluding ties, Voxtral is approximately balanced against ElevenLabs v3 under explicit steering and wins more often against both ElevenLabs systems under implicit steering. Its 35.4% and 37.1% win rates against Gemini show that the flagship-voice results do not establish overall human-preference leadership.
4.3.2 Zero-Shot Voice Cloning. The evaluation uses high-quality recordings from two recognized speakers per language and asks listeners to judge voice likeness, naturalness, and expressivity. Against ElevenLabs Flash v2.5, Voxtral TTS achieves a 68.4% overall micro-average win rate.
- Table 5 reports Arabic 72.9%, Dutch 49.4%, English 60.8%, French 54.4%, German 72.0%, Hindi 79.8%, Italian 57.1%, Portuguese 74.4%, and Spanish 87.8%.
- Eight languages exceed 50%; Dutch is slightly below it. Without uncertainty estimates, the Dutch result establishes neither a strict win nor statistical equivalence.
- The higher aggregate cloning win rate than the flagship-voice win rate shows stronger relative performance in this cloning setup, but the different evaluation conditions do not establish unrestricted generalization to user recordings.
The overall 68.4% cloning win rate combines substantial language variation, from 49.4% for Dutch to 87.8% for Spanish. Eight of nine reported language rates exceed 50%; without uncertainty estimates, the Dutch value supports neither a strict win nor a demonstrated equivalence.
5 Analysis
The analysis compares pretrained and DPO checkpoints, then varies the acoustic solver’s number of function evaluations and CFG scale. Automatic metrics quantify these changes, while qualitative observations address hallucinations, volume consistency, and responsiveness to emotions implicit in text.
5.1 DPO improvements
Table 6 shows that DPO lowers WER on eight MiniMax-TTS languages and SEED-TTS while increasing UTMOS on every reported row. The largest WER improvements are German, from 4.08% to 0.83%, and French, from 5.01% to 3.22%; Hindi worsens from 3.39% to 4.99%.
- Hindi UTMOS nevertheless increases from 3.43 to 3.56, illustrating that predicted naturalness and intelligibility can move in different directions.
- The authors qualitatively report fewer hallucinations, fewer skipped words, and less volume tapering after DPO.
- Speaker similarity changes remain within ±0.01 of the pretrained checkpoint, according to the text; those scores are not tabulated.
DPO improves WER in every reported task except Hindi and raises UTMOS throughout. German WER falls from 4.08% to 0.83%, whereas Hindi rises from 3.39% to 4.99% despite a UTMOS improvement, exposing a language-specific trade-off in preference optimization.
5.2 Inference Parameters
Figure 4 supports 8 NFEs as a quality–cost compromise: UTMOS rises from 1.44 at 2 NFEs to 3.48 at 8, and speaker similarity rises from 0.421 to 0.732. Increasing to 16 leaves UTMOS at 3.48, raises similarity only to 0.735, and worsens WER from 6.7 to 7.0.
- WER improves sharply from 30.1 at 2 NFEs to 6.5 at 4, but is 6.7 at 8; additional evaluations do not improve it monotonically.
- Increasing CFG α from 1.0 to 1.4 reduces plotted WER from 7.5 to 6.2 and raises speaker similarity from 0.707 to 0.737. UTMOS peaks at 3.52 for α=1.3 before falling to 3.48 at α=1.4.
- Internal listening found excessive reference adherence at higher guidance, reducing responsiveness to emotions implicit in text. The default α=1.2 favors high-quality recordings; the paper suggests that in-the-wild recordings may benefit from stronger guidance.
The NFE plots show strong gains in UTMOS and speaker similarity through 8 evaluations, followed by diminishing returns; WER is lowest at 4 evaluations rather than 8. Higher CFG improves plotted WER and similarity but not monotonically UTMOS, while internal listening motivates the lower default guidance to preserve text-implied emotional variation.
6 Inference and Serving in vLLM-Omni
vLLM-Omni serves Voxtral TTS as a two-stage pipeline: semantic/acoustic token generation followed by codec waveform decoding. Shared-memory asynchronous chunk transfer allows playback to begin before the complete waveform is synthesized.
6.1 CUDA Graph Acceleration for Flow-Matching Transformer
The flow-matching transformer is the reported generation-stage computational bottleneck because N integration evaluations with CFG require 2 × N forward passes per frame. Capturing the entire ODE solver in CUDA graphs reduces Python overhead and kernel-launch latency.
- Startup warms up and captures graphs for batch-size buckets. Inference pads to the nearest bucket, replays its graph, and slices outputs back to the actual batch size; batches above the largest bucket fall back to eager execution.
- For a 500-character input, 10-second reference, concurrency 1, and one H200, Table 7 reports latency dropping from 133 ms to 70 ms and RTF from 0.258 to 0.103.
- The paper describes these as a 47% latency improvement and a 2.5× reduction in RTF.
For concurrency 1 on one H200 with a 500-character input and 10-second reference, CUDA graph execution lowers latency from 133 ms to 70 ms and RTF from 0.258 to 0.103. This measures an implementation benefit from capturing the repeated acoustic solver under the stated serving conditions.
6.2 Asynchronous Chunked Streaming
Generation and codec decoding use separate scheduling loops. A transfer manager accumulates per-request audio tokens and emits a chunk once a predefined buffer length is reached, allowing autoregressive generation to overlap waveform synthesis.
- Each chunk includes preceding frames as well as new frames, providing temporal context for the codec decoder’s causal sliding-window attention across chunk boundaries.
- The PDF does not specify the chunk length or overlap length.
6.3 Inference Throughput
Table 8 evaluates one NVIDIA H200 at concurrency 1, 16, and 32 with 500-character text inputs and 10-second voice references. Throughput increases from 119.14 to 879.11 to 1430.78 char/s/GPU, while time to first audio increases from 70 to 331 to 552 ms and RTF from 0.103 to 0.237 to 0.302.
- The client wait rate—the fraction of audio chunks requiring the client to stall—remains 0% at all three concurrency levels.
- These measurements demonstrate uninterrupted real-time streaming for the tested workload, with approximately 12× higher aggregate throughput at concurrency 32 than at concurrency 1.
- The reported conditions do not test longer inputs, different reference lengths, heterogeneous request distributions, or other GPUs.
Increasing concurrency from 1 to 32 raises throughput from 119.14 to 1430.78 char/s/GPU while time to first audio remains below one second and wait rate remains 0%. The per-request RTF rises to 0.302, indicating continued real-time operation for this fixed workload rather than unchanged per-request speed.
7 Conclusion
The paper concludes that combining autoregressive semantic generation, flow-matched acoustic generation, and an ASR-distilled VQ–FSQ codec enables expressive multilingual voice cloning from short references. Voxtral TTS is released as open weights under CC BY-NC, rather than an unrestricted commercial-use license.
- Webpage: https://mistral.ai/news/voxtral-tts
- Model weights: https://huggingface.co/mistralai/Voxtral-4B-TTS-2603
Appendix
- The 16-page PDF contains no appendix. It identifies the paper as arXiv:2603.25551v2 [cs.AI], dated 6 Apr 2026, and ends with contributor lists, acknowledgements, and references. An implementation-relevant notation inconsistency remains: Section 2.3 defines x0 as Gaussian noise and x1 as acoustic data, whereas Equations (6)–(7) and their explanation define x0 as data and x1 as Gaussian noise.
Brief Thoughts
The main methodological contribution is the alignment between the codec representation and the generator: discrete semantic prediction retains a low-rate temporal interface, while continuous acoustic generation avoids 36 depth-wise autoregressive steps. The evidence supports competitive voice cloning and low-latency streaming under the tested conditions, but does not isolate the contribution of every architectural choice or establish performance across deployment workloads. Human preference complements automatic evaluation: speaker similarity consistently favors Voxtral, yet WER varies and Gemini wins the flagship-voice comparisons. Explicit steering also uses different control interfaces, so it compares complete systems rather than identical conditioning mechanisms. The PDF does not report total training-data volume, detailed data composition, human-preference confidence intervals, or numerical results for the alternative acoustic architectures and codec ablations. These omissions limit reproducibility and uncertainty assessment. Cloning evaluation uses two recognized speakers per language and high-quality reference recordings. It does not quantify generalization to noisy microphones, larger speaker populations, or cross-language reference–target combinations. DPO improves most intelligibility results but degrades Hindi WER, and speaker similarity barely changes. Its benefits are metric- and language-dependent rather than universal.