arXiv:2510.18724cs.CLcs.LG2025-10被引 2

通过标注语言切换点提升模型在混用语言语音中的表现。

Adapting Language Balance in Code-Switching Speech

  • 用可微分代理信号标记语言切换位置,增强模型关注
  • 阿拉伯语和中英混用数据上替换错误率下降
  • 适合研究多语言语音识别与生成的学者

尽管大型基础模型在标准基准测试中表现优异,但在处理语言混用场景时仍表现不佳。当数据稀缺无法解释性能差时,问题可能源于语言切换时刻出现频率低,且第二语言嵌入信号较隐蔽。我们不依赖模型自行学习这种稀有性,而是主动在训练中引入标注。为准确评估性能,需精确定位语言切换点,因这些位置的识别错误影响最大。基于嵌入语言与主语言的差异,我们设计了一种可微分的代理信号来突出这些切换点,从而强化模型在关键位置的学习。该方法有效缓解了生成过程中的上下文偏差——语言混用的核心挑战,提升了模型鲁棒性。在阿拉伯语和中英混用数据上的实验表明,模型能更准确预测切换位置,替换错误率显著降低。

原文摘要 · Abstract (English)

Despite achieving impressive results on standard benchmarks, large foundational models still struggle against code-switching test cases. When data scarcity cannot be used as the usual justification for poor performance, the reason may lie in the infrequent occurrence of code-switched moments, where the embedding of the second language appears subtly. Instead of expecting the models to learn this infrequency on their own, it might be beneficial to provide the training process with labels. Evaluating model performance on code-switching data requires careful localization of code-switching points where recognition errors are most consequential, so that the analysis emphasizes mistakes occurring at those moments. Building on this observation, we leverage the difference between the embedded and the main language to highlight those code-switching points and thereby emphasize learning at those locations. This simple yet effective differentiable surrogate mitigates context bias during generation -- the central challenge in code-switching -- thereby improving the model's robustness. Our experiments with Arabic and Chinese-English showed that the models are able to predict the switching places more correctly, reflected by the reduced substitution error.

语音识别语言混用多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。