arXiv:2606.19381cs.SDcs.AI2026-06中稿 · Interspeech 2026被引 1

用语言切换指数指导合成语音,提升双语语音识别效果。

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

论文配图:Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
图 1 · 摘自论文原文
  • 引入语言切换指数引导合成语音生成
  • 在测试集上将混合错误率降低至8.9%和14.2%
  • 适合需要双语语音识别增强的研究者

代码切换(CS)自动语音识别(ASR)因高质量双语文本-语音配对数据稀缺而面临挑战。尽管已有研究尝试通过文本转语音(TTS)生成合成数据,但现有方法主要优化语音重建保真度,未显式保证语言边界一致性,限制了其在CS ASR中的增广效果。本文提出一种基于代码混合引导的偏好学习框架,利用代码混合指数(CMI)引导合成语音生成,提升代码切换保真度。在SEAME普通话-英语对话语料库上的实验表明,该方法显著增强了合成数据在ASR微调中的实用性。具体而言,在微调Whisper Large模型时,开发集DevMAN和DevSGE的混合错误率(MER)分别从12.1%/17.8%降至8.9%/14.2%。

原文摘要 · Abstract (English)

Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored, existing CS TTS approaches primarily optimise reconstruction fidelity and do not explicitly enforce language-boundary consistency, thereby limiting their effectiveness for CS ASR augmentation. This paper proposes a code-mixing guided preference-learning framework that steers synthetic speech generation toward improved code-switching fidelity using the Code Mixing Index (CMI). Experiments on the SEAME Mandarin-English conversational corpus demonstrate that the proposed method enhances the utility of synthetic data for ASR fine-tuning. Specifically, when fine-tuning Whisper Large, the proposed approach reduces Mixed Error Rate (MER) from 12.1%/17.8% to 8.9%/14.2% on the DevMAN and DevSGE sets, respectively.

语音识别双语语音合成数据TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。