arXiv:2606.21990cs.CLeess.AS2026-06中稿 · INTERSPEECH 2026

用少量合成数据提升多语种语音识别的混语处理能力。

Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR

  • 通过贝叶斯因子化适配,高效融合混语知识而不破坏原有性能。
  • 在混语词上错误率降低32.87%,整体词错率改善5.31%。
  • 适合需要高鲁棒性混语识别的实用语音系统部署场景。

真实场景中,混语(Code-switching, CSW)仍是大型多语种自动语音识别系统的主要挑战。尽管可通过合成混语数据微调,但通常会损害原有强单语基线性能。本文目标是在保持强单语能力的同时,扩展模型以应对复杂混语,包括跨语言形态变化。提出贝叶斯因子化适配方法,无需覆盖已有知识即可高效集成混语相关能力。仅需少量合成数据,该方法在混语词上将错误率降低32.87%,整体词错率(WER)改善5.31%,同时维持单语性能。结果表明,有效混语适应更依赖知识整合而非数据复杂度。

原文摘要 · Abstract (English)

Code-switching (CSW) remains challenging for large multi-lingual ASR systems in real-world deployment. While fine-tuning on synthetic CSW data is possible, it generally degrades strong monolingual baselines. Our goal is to preserve these capabilities while extending models to handle complex code-switching, including morphological variations across languages. We propose Bayesian factorized adaptation, which learns to efficiently integrate switching-relevant knowledge into strong pretrained models without overwriting existing capabilities. Requiring only a small amount of synthetic data, our approach reduces transcription errors by 32.87% on code-switched words while improving overall WER by 5.31%, all while maintaining mono-lingual performance. Our results demonstrate that effective CSW adaptation depends more on knowledge integration than data complexity.

语音识别混语处理多语种贝叶斯适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。