arXiv.org
Voxtral TTS
Voxtral TTS is a multilingual zero-shot text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. It uses a hybrid architecture: an autoregressive decoder backbone (based on Ministral 3B) predicts semantic speech tokens, while a flow-matching transformer predicts acoustic tokens. The tokens are produced by Voxtral…
Mistral-AI, :, Alexander H. Liu, Alexis Tacnet, et al.- Published
- Mar 2026
- Citations
- 0
- Code
- Not linked