arXiv:2607.23938eess.AS2026-07被引 3

Qwen-Audio-3.0-TTS实现高鲁棒、多语言、可自由控制的语音合成。

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

论文配图:Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
图 1 · 摘自论文原文
  • 分阶段训练+低帧率语音分词,降低延迟提升效率。
  • 支持16种语言、20个中文方言,最长3分钟单次生成。
  • 可通过自然语言指令精细控制,适合实际应用部署。

本文介绍 Qwen-Audio-3.0-TTS,一个面向生产环境的语音合成系统,同时提升内容一致性、说话人相似性、韵律自然度、音频质量、可控性、多语言覆盖、效率和鲁棒性。系统采用12.5 Hz低帧率语音分词器以减少推理延迟,并引入五阶段渐进式训练范式,协同优化语言模型(LM)与流匹配模型(FM)。通过自由风格的自然语言指令和细粒度内联标签实现生产级控制,支持16种语言、20个中文方言区域,单次生成最长可达3分钟,并能从噪声、混响或模糊的参考语音中稳健生成。在SEED-TTS-Eval、CV3-Eval、指令遵循、长文本合成及声学鲁棒性评测中,该模型在多个维度达到当前最佳表现或综合得分最高,且在独立的人工分析语音合成排行榜上排名第一。这些结果确立了 Qwen-Audio-3.0-TTS 作为生产级语音合成的坚实基础。

原文摘要 · Abstract (English)

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated language model (LM) and flow-matching model (FM) optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.

语音合成多语言可控生成生产级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。