arXiv:2412.08651eess.AScs.CL2024-12

通过语言边界对齐提升多语混用语音识别效果

Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection

  • 在编码器中间层注入语言识别信息,增强语言表征
  • 引入非尖峰CTC损失与语言边界对齐,提升多语切换识别精度
  • 适合需要高精度多语混用识别的语音系统研发者

多语混用——多语言使用者在对话中交替使用不同语言——仍给端到端(E2E)语音识别系统带来显著挑战,源于声学与语义混淆。现有系统难以有效应对语言快速切换,导致性能严重下降。本文提出三项创新:首先,在编码器多个中间层融合语言识别(LID)信息,使输出表征包含更细致的语言特征;其次,通过新颖的语言边界对齐损失,使后续ASR模块能更有效地利用内部语言后验知识;第三,探索利用语言后验实现共享编码器与语言特定编码器间的深度交互。在SEAME数据集上的全面实验表明,该方法优于先前最优的基于解耦混合专家(D-MoE)的方法,进一步提升了编码器对语言的敏感度。

原文摘要 · Abstract (English)

Code-switching-where multilingual speakers alternately switch between languages during conversations-still poses significant challenges to end-to-end (E2E) automatic speech recognition (ASR) systems due to phenomena of both acoustic and semantic confusion. This issue arises because ASR systems struggle to handle the rapid alternation of languages effectively, which often leads to significant performance degradation. Our main contributions are at least threefold: First, we incorporate language identification (LID) information into several intermediate layers of the encoder, aiming to enrich output embeddings with more detailed language information. Secondly, through the novel application of language boundary alignment loss, the subsequent ASR modules are enabled to more effectively utilize the knowledge of internal language posteriors. Third, we explore the feasibility of using language posteriors to facilitate deep interaction between shared encoder and language-specific encoders. Through comprehensive experiments on the SEAME corpus, we have verified that our proposed method outperforms the prior-art method, disentangle based mixture-of-experts (D-MoE), further enhancing the acuity of the encoder to languages.

语音识别多语混用语言识别深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。