arXiv:2505.07202cs.CLcs.SD2025-05

对比两种对话语音训练方法,发现按句训练更高效且音质更好。

On the Cost and Benefits of Training Context with Utterance or Full Conversation Training: A Comparative Stud

  • 用上下文条件训练单句语音,提升生成质量
  • 单句训练得分4.3分(满分5分),快37%且无说话人混淆
  • 适合想高效开发高质量对话语音系统的人

当前对话式语音合成系统虽能生成高质量语音,但大多未开源。究竟是架构不足,还是训练方法有问题?本文通过20个GPU小时的实验证明:基于上下文的单句训练在语音质量(MOS 4.3/5.0)上优于完整对话训练(3.7/5.0),同时节省37%训练时间,且避免说话人相似性幻觉问题。研究为对话语音合成提供了高效实用的训练策略,推荐采用带上下文条件的单句训练方式。

原文摘要 · Abstract (English)

Modern TTS systems designed for conversations achieve high-quality utterances but often remain inaccessible publicly. Are existing open-source architectures inadequate, or are current training techniques insufficient? This paper investigates prominent models and their underlying behaviors regarding conversational context. Using 20 GPU-hours on an NVIDIA H100, we empirically examine two approaches: context-based utterance-level training versus full conversation training. Results demonstrate that context-based utterance training achieves superior MOS scores (4.3/5.0 vs 3.7/5.0) and reduces training time by 37%, while full conversation approaches suffer from speaker similarity hallucination issues. These findings provide practical guidelines for conversational TTS development, favoring utterance-level training with contextual conditioning for both resource efficiency and output quality.

语音合成训练效率对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。