arXiv:2601.00935eess.AScs.AI2026-01中稿 · APSIPA 2025被引 2

用多语言语音合成生成数据,提升中英混杂语音识别准确率

Improving Code-Switching Speech Recognition with TTS Data Augmentation

  • 用CosyVoice2模型在SEAME数据集上训练,生成中英混杂语音
  • 合成数据使开发集错误率降低1.8%~1.9%,效果稳定
  • 适合低资源混杂语种语音识别研究者参考

对话式中英混杂语音识别因真实高质量标注数据稀缺而面临挑战。本文探索使用多语言文本转语音(TTS)模型作为数据增强手段。具体而言,我们基于SEAME数据集微调多语言CosyVoice2 TTS模型,生成合成的中英混杂对话语音,显著增加了训练数据的数量与说话人多样性。实验表明,将真实语音与合成语音结合后,DevMan测试集的混合错误率(MER)从12.1%降至10.1%,DevSGE测试集从17.8%降至16.0%,性能持续提升。结果验证了多语言TTS是提升低资源混杂语种语音识别鲁棒性的有效且实用工具。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explores multilingual text-to-speech (TTS) models as an effective data augmentation technique to address this shortage. Specifically, we fine-tune the multilingual CosyVoice2 TTS model on the SEAME dataset to generate synthetic conversational Chinese-English code-switching speech, significantly increasing the quantity and speaker diversity of available training data. Our experiments demonstrate that augmenting real speech with synthetic speech reduces the mixed error rate (MER) from 12.1 percent to 10.1 percent on DevMan and from 17.8 percent to 16.0 percent on DevSGE, indicating consistent performance gains. These results confirm that multilingual TTS is an effective and practical tool for enhancing ASR robustness in low-resource conversational code-switching scenarios.

语音识别数据增强混杂语种TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。