AdaCS通过自适应注意力机制,显著提升越南语跨语段语音识别准确率。
AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
- 引入自适应偏置注意力模块,动态识别并归一化跨语段内容
- 在两个新测试集上分别降低56.2%和36.8%的词错误率
- 适用于低资源语言,对未见领域跨语段现象有强适应性
句内跨语段(CS)指单个话语中语言交替的现象,是自动语音识别(ASR)系统的重要挑战。例如,越南语使用者在表达中嵌入外语专有名词或术语。现有ASR模型因训练数据多为单语种且跨语段模式不可预测,难以准确转录此类内容,尤其在低资源语言场景下更为突出。本文提出AdaCS,一种在编码器-解码器架构中集成自适应偏置注意力模块(BAM)的归一化模型。该方法利用推理时提供的偏置词表,动态识别并归一化跨语段短语,显著增强模型对未见领域的适应能力。实验表明,AdaCS在两个新构建的越南语跨语段语音识别测试集上,相比此前最优方法分别实现56.2%和36.8%的词错误率(WER)下降,验证了其有效性与泛化能力。
原文摘要 · Abstract (English)
Intra-sentential code-switching (CS) refers to the alternation between languages that happens within a single utterance and is a significant challenge for Automatic Speech Recognition (ASR) systems. For example, when a Vietnamese speaker uses foreign proper names or specialized terms within their speech. ASR systems often struggle to accurately transcribe intra-sentential CS due to their training on monolingual data and the unpredictable nature of CS. This issue is even more pronounced for low-resource languages, where limited data availability hinders the development of robust models. In this study, we propose AdaCS, a normalization model integrates an adaptive bias attention module (BAM) into encoder-decoder network. This novel approach provides a robust solution to CS ASR in unseen domains, thereby significantly enhancing our contribution to the field. By utilizing BAM to both identify and normalize CS phrases, AdaCS enhances its adaptive capabilities with a biased list of words provided during inference. Our method demonstrates impressive performance and the ability to handle unseen CS phrases across various domains. Experiments show that AdaCS outperforms previous state-of-the-art method on Vietnamese CS ASR normalization by considerable WER reduction of 56.2% and 36.8% on the two proposed test sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。