开源100万小时语音数据和模型,让开放TTS达到顶尖水平
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech

- 用公开数据训练扩散变换器模型,覆盖510万小时高质量语音
- 10亿参数模型在多项评测中媲美闭源大模型,尤其在语音相似度上领先
- 提供完整评估基准,适合研究者复现与测试语音鲁棒性
近期文本转语音(TTS)模型在语音自然度和质量上取得显著进展,但大规模开放数据在其中的作用仍不明确。本文提出Raon-OpenTTS,一个性能可媲美先进闭源模型的开源TTS系统,以及其基础数据集Raon-OpenTTS-Pool,包含615万小时、2.4亿段英语语音片段,来自公开语料库与网络采集。通过模型驱动的过滤流程,从中筛选出510万小时、1.94亿段的高质量子集Raon-OpenTTS-Core。基于此,我们训练了从0.3B到1B参数的扩散变换器(DiT)系列模型。在多个基准测试中,Raon-OpenTTS-1B的表现与采用数百万小时专有数据训练的Qwen3-TTS和CosyVoice 3相当。在Seed-TTS-Eval上,其词错误率(WER)为1.78%,说话人相似度(SIM)达0.749,分别位列第二和第一;在CV3-Hard-EN上,WER为6.15%,SIM为0.775,两项指标均排名第一。此外,我们构建了结构化评估基准Raon-OpenTTS-Eval,用于评测不同声学条件下(清晰、嘈杂、真实场景、情感表达)的鲁棒性。在此基准上,该模型平均WER和SIM最优,人类偏好得分(CMOS)仅次于最佳模型。相关数据、代码与模型权重均已开源。
原文摘要 · Abstract (English)
Recent advances in text-to-speech (TTS) models show impressive speech naturalness and quality, yet the role of large-scale open data in driving this progress remains underexplored. In this work, we introduce Raon-OpenTTS, an open TTS model that performs competitively with state-of-the-art closed-data TTS models, and Raon-OpenTTS-Pool, a large-scale open dataset for reproducible TTS training. Raon-OpenTTS-Pool consists of 615K hours of 240M speech segments aggregated from publicly available English speech corpora and web-sourced recordings. With a model-based filtering pipeline applied to Raon-OpenTTS-Pool, we derive Raon-OpenTTS-Core, a curated, high-quality subset of 510K hours and 194M speech segments. Using Raon-OpenTTS-Core, we train Raon-OpenTTS, a series of diffusion transformer (DiT)-based TTS models from 0.3B to 1B parameters. On multiple benchmarks, Raon-OpenTTS-1B shows comparable performance to state-of-the-art models such as Qwen3-TTS and CosyVoice 3, which are trained on several million hours of proprietary speech data. Notably, on Seed-TTS-Eval, Raon-OpenTTS-1B achieves a word error rate (WER) of 1.78% and a speaker similarity (SIM) of 0.749, ranking second on WER and first on SIM among recent open-weight TTS baselines. On CV3-Hard-EN, Raon-OpenTTS-1B achieves a WER of 6.15% and a SIM of 0.775, ranking first on both metrics. Furthermore, to support robust evaluation, we introduce Raon-OpenTTS-Eval, a structured benchmark for assessing TTS robustness across diverse acoustic conditions including clean, noisy, in-the-wild, and expressive speech. On Raon-OpenTTS-Eval, Raon-OpenTTS-1B achieves the best average WER and SIM among all evaluated models, and the second-best human preference, as measured by comparative mean opinion score (CMOS). Our data pool, filtering pipeline, training code, and checkpoints are publicly available at https://github.com/krafton-ai/RAON-OpenTTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。