仅需3秒参考音频,即可生成自然且富有表现力的多语言语音。
Voxtral TTS
- 混合架构:先生成语义标记,再用流匹配生成声学标记。
- 在母语者评测中,68.4%胜过ElevenLabs Flash v2.5。
- 适合语音克隆场景,支持多语言,开源可商用(CC BY-NC)
我们提出Voxtral TTS,一种能从仅3秒参考音频生成自然语音的多语言文本到语音模型。该模型采用混合架构,先自回归生成语义语音标记,再通过流匹配生成声学标记。所有标记由Voxtral Codec编码解码,该编码器基于混合VQ-FSQ量化方案从零训练。在母语者参与的人类评估中,Voxtral TTS因自然度和表现力,在语音克隆任务中获得68.4%的胜率,优于ElevenLabs Flash v2.5。模型权重已发布,采用CC BY-NC许可。
原文摘要 · Abstract (English)
We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid architecture that combines auto-regressive generation of semantic speech tokens with flow-matching for acoustic tokens. These tokens are encoded and decoded with Voxtral Codec, a speech tokenizer trained from scratch with a hybrid VQ-FSQ quantization scheme. In human evaluations conducted by native speakers, Voxtral TTS is preferred for multilingual voice cloning due to its naturalness and expressivity, achieving a 68.4\% win rate over ElevenLabs Flash v2.5. We release the model weights under a CC BY-NC license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。