80亿参数语音合成模型,自然度与克隆精度均达顶尖水平。
ZONOS2 Technical Report

- 采用新型专家混合架构,模型规模扩至80亿参数,推理更快更高效。
- 训练数据扩充至600万小时,显著提升语音自然度和声线还原精度。
- 开源模型权重与代码,适合语音生成、语音克隆方向的研究者使用。
我们提出ZONOS2 8B,最新一代语音合成模型,在自然度、语调表现及声线克隆保真度上达到业界领先水平。相比Zonos-v0.1,我们在模型规模、训练数据和训练方法上均实现优化:将模型参数从16亿增至80亿(活跃参数9亿),采用新型专家混合(MoE)结构,提升推理延迟与吞吐效率;训练语料库从20万小时扩展至超过600万小时,依赖新数据处理流水线;简化后训练与条件控制流程,进一步提升语音自然度与声线克隆准确性。我们在质量、说话人相似性、词错误率(WER)及ZTTS1-Eval这一新型语音合成评估基准上进行评测,结果表明其性能可媲美当前最先进系统,同时保持良好流式处理延迟。模型权重与示例推理代码已通过Apache 2.0许可证在GitHub与Hugging Face公开。
原文摘要 · Abstract (English)
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across scale, data, and training recipe. We scale the model from 1.6B to 8B total parameters (900M active) with a novel mixture-of-experts (MoE) backbone, improving inference latency and throughput. We expand our training corpus from 200K to over 6M hours using a new data processing pipeline, and we simplify our post-training and conditioning recipes to improve naturalness and voice cloning fidelity. We evaluate ZONOS2 8B on quality, speaker similarity, WER, and ZTTS1-Eval, our novel TTS benchmark, where it performs competitively with state-of-the-art systems while maintaining good streaming latency. We release our model weights and example inference code under an Apache 2.0 license on GitHub and Hugging Face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。