arXiv:2607.02862cs.CLeess.AS2026-07被引 2

通过融合语音与文本特征,同时提升印度语种的语音识别与方言识别效果。

Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion

论文配图:Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion
图 1 · 摘自论文原文
  • 用瓶颈编码器提取语音方言特征,结合RoBERTa处理识别结果嵌入
  • 在33种方言上实现81.63%的方言识别准确率,语音识别词错误率17.73%
  • 适合低资源印度语言的联合建模,对语音与方言研究者有参考价值

自动语音识别(ASR)和方言识别(DID)对印度语言至关重要,许多语言属于低资源类型且方言差异显著。现有方法通常单独优化ASR或DID,导致性能权衡。本文提出一种多模态框架,实现ASR与DID的联合优化。该方法采用瓶颈编码器从基于Conformer的语音表示中提取方言特征,并使用RoBERTa编码器处理ASR生成的CTC嵌入。通过门控机制融合两类特征,再经注意力编码器优化表示。最终将学习到的嵌入与Conformer输出拼接,增强语音识别特征。在涵盖八种印度语言、三十三种方言的数据集上评估,平均DID准确率达81.63%,平均词错误率(CER)为4.65%,字错误率(WER)为17.73%。结果验证了该方法在联合建模中的有效性。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences. Existing methods often optimize ASR or DID individually, resulting in performance trade-offs. In this work, we propose a multimodal framework that jointly improves ASR and DID. Our method employs a Bottleneck Encoder to extract dialectal features from Conformer-based speech representations and a RoBERTa encoder to process ASR-generated CTC embeddings. A gating mechanism merges these features, followed by an attention encoder to refine the representations. The learned embeddings are concatenated with Conformer outputs to enhance ASR features. Evaluated on eight Indian languages with thirty-three dialects, our method achieves an average DID accuracy of 81.63% and average CER and WER of 4.65% and 17.73%, respectively. These results highlight the effectiveness of our method for joint ASR-DID modeling.

语音识别方言识别多模态印度语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。