arXiv:2512.17356cs.SD2025-12被引 1

纯合成数据可训练出优于真实数据的语音合成模型

Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization Capability

  • 用纯合成数据训练语音模型,通过控制文本与说话人多样性提升性能
  • 合成数据训练的模型在清晰度和鲁棒性上超越真实数据训练模型
  • 适合追求高质量语音生成且无真实录音场景的研究者

合成数据在语音合成(TTS)训练中的潜力日益受到关注,但其合理性与有效性仍需系统验证。本研究系统考察了仅使用合成数据进行TTS训练的可行性,并分析了文本丰富度、说话人多样性、噪声水平及语调风格等因素对模型表现的影响。实验表明,增加说话人和文本多样性显著提升合成质量与鲁棒性;更清洁的训练数据(低噪声)进一步改善性能。此外,标准语调风格有助于模型更高效学习。实验显示,在相似条件下,纯合成数据训练的模型具备超越真实数据训练模型的潜力,原因在于避免了真实世界中的不完美与噪声。

原文摘要 · Abstract (English)

The potential of synthetic data in text-to-speech (TTS) model training has gained increasing attention, yet its rationality and effectiveness require systematic validation. In this study, we systematically investigate the feasibility of using purely synthetic data for TTS training and explore how various factors--including text richness, speaker diversity, noise levels, and speaking styles--affect model performance. Our experiments reveal that increasing speaker and text diversity significantly enhances synthesis quality and robustness. Cleaner training data with minimal noise further improves performance. Moreover, we find that standard speaking styles facilitate more effective model learning. Our experiments indicate that models trained on synthetic data have great potential to outperform those trained on real data under similar conditions, due to the absence of real-world imperfections and noise.

语音合成合成数据TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。