用语言切换指数指导合成语音,提升双语语音识别效果。
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

- 引入语言切换指数引导合成语音生成
- 在测试集上将混合错误率降低至8.9%和14.2%
- 适合需要双语语音识别增强的研究者
代码切换(CS)自动语音识别(ASR)因高质量双语文本-语音配对数据稀缺而面临挑战。尽管已有研究尝试通过文本转语音(TTS)生成合成数据,但现有方法主要优化语音重建保真度,未显式保证语言边界一致性,限制了其在CS ASR中的增广效果。本文提出一种基于代码混合引导的偏好学习框架,利用代码混合指数(CMI)引导合成语音生成,提升代码切换保真度。在SEAME普通话-英语对话语料库上的实验表明,该方法显著增强了合成数据在ASR微调中的实用性。具体而言,在微调Whisper Large模型时,开发集DevMAN和DevSGE的混合错误率(MER)分别从12.1%/17.8%降至8.9%/14.2%。
原文摘要 · Abstract (English)
Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored, existing CS TTS approaches primarily optimise reconstruction fidelity and do not explicitly enforce language-boundary consistency, thereby limiting their effectiveness for CS ASR augmentation. This paper proposes a code-mixing guided preference-learning framework that steers synthetic speech generation toward improved code-switching fidelity using the Code Mixing Index (CMI). Experiments on the SEAME Mandarin-English conversational corpus demonstrate that the proposed method enhances the utility of synthetic data for ASR fine-tuning. Specifically, when fine-tuning Whisper Large, the proposed approach reduces Mixed Error Rate (MER) from 12.1%/17.8% to 8.9%/14.2% on the DevMAN and DevSGE sets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。