arXiv:2606.05846cs.CLeess.AS2026-06

让语音识别模型跨语言切换,通用性有限但有潜力。

Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

论文配图:Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs
图 1 · 摘自论文原文
  • 用模型融合和领域泛化方法尝试跨语言对迁移能力
  • 仅在少量已见语言对上训练的模型可微弱泛化到新语言对
  • 适合研究多语言语音识别泛化性的学者参考

自动语音识别(ASR)已成为人机交互的关键技术。然而,由于多样语言对之间的多语种混用语音资源严重匮乏,混用语言语音识别(CS-ASR)仍极具挑战。现有方法主要通过合成混用语音或在有限双语数据集上进行特定语言对微调来提升性能。然而,这些方法存在固有的可扩展性限制:每新增一个语言对,都需单独构建支持,其数量随语言数呈组合爆炸式增长。本文探讨是否可通过模型融合与领域泛化方法,将从有限已见语言对中学到的混用能力泛化至未见过的语言对。实验表明,合并后的双语混用模型可实现对未见语言对的微弱泛化,提示双语混用能力在不同语言对间转移能力有限。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) has become a key technology for human--AI interaction. However, code-switching ASR (CS-ASR) remains particularly challenging due to the severe scarcity of multilingual CS speech resources across diverse language pairs. Existing approaches primarily improve CS-ASR performance through synthetic CS speech generation or pair-specific fine-tuning on limited bilingual datasets. Nevertheless, these approaches face an inherent scalability limitation, as support for CS must be developed separately for language pairs whose number grows combinatorially with the number of supported languages. In this work, we investigate whether CS capabilities learned from a limited set of seen language pairs can generalize to unseen language pairs through model merging and domain generalization methods. Our experiments show that merged bilingual CS-ASR models modestly generalize to unseen language pairs, suggesting limited transfer of bilingual CS capabilities across language pairs.

语音识别多语言模型泛化代码切换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。