arXiv:2503.07522eess.AScs.CL2025-03

让英语语音识别模型同时听懂印地语,且不影响英语识别效果。

Building English ASR model with regional language support

  • 用共享隐藏层+语言特异性投影层,通过自注意力动态加权
  • 印地语字错误率降69.3%,英语降5.7%(相比单语英语模型)
  • 适合需要中英文混识或印地语语音交互的场景

本文提出一种新型英语自动语音识别(ASR)系统,可有效处理印地语语音输入,且不降低英语识别性能。我们设计了一种名为SplitHead with Attention(SHA)的声学模型,其包含跨语言共享的隐层和语言特异的投影层,并通过自注意力机制动态计算各语言权重,按输入数据自动分配对应投影层贡献。此外,提出一种语言建模方法,融合英语与音译印地语语料的n-gram模型进行插值。实验结果表明,与单语英语模型相比,该方法在印地语测试集上实现69.3%的相对字错误率下降,在英语测试集上实现5.7%的相对下降。

原文摘要 · Abstract (English)

In this paper, we present a novel approach to developing an English Automatic Speech Recognition (ASR) system that can effectively handle Hindi queries, without compromising its performance on English. We propose a novel acoustic model (AM), referred to as SplitHead with Attention (SHA) model, features shared hidden layers across languages and language-specific projection layers combined via a self-attention mechanism. This mechanism estimates the weight for each language based on input data and weighs the corresponding language-specific projection layers accordingly. Additionally, we propose a language modeling approach that interpolates n-gram models from both English and transliterated Hindi text corpora. Our results demonstrate the effectiveness of our approach, with a 69.3% and 5.7% relative reduction in word error rate on Hindi and English test sets respectively when compared to a monolingual English model.

语音识别多语言自注意力印地语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。