构建多语言语音识别系统,提升印地语系8语言33方言的识别准确率。
Building Robust and Scalable Multilingual ASR for Indian Languages
- 采用多解码器架构与音素公共标签集作为中间表示。
- 在3种语言上达到最低词错误率,语言与方言识别准确率最高。
- 适合需要高鲁棒性多语言语音识别的研究者与开发者。
本文介绍由印度理工学院马德拉斯分校SPRING实验室为ASRU MADASR 2.0挑战赛开发的系统。该系统聚焦于在8种语言、33种方言中提升语音识别对语言和方言的判别能力。团队参与了需从零构建多语言系统的第1、第2赛道,限制使用额外数据。提出一种新型训练方法:基于多解码器架构与音素公共标签集(CLS)作为中间表示,显著提升基线模型在CLS空间的表现。同时探讨多种方法,在将音素表示还原为对应音形表示时保留性能增益。最终系统在第2赛道中,3种语言的词错误率(WER)与字符错误率(CER)优于基线,并在所有参赛队伍中取得最高的语言识别与方言识别准确率。
原文摘要 · Abstract (English)
This paper describes the systems developed by SPRING Lab, Indian Institute of Technology Madras, for the ASRU MADASR 2.0 challenge. The systems developed focuses on adapting ASR systems to improve in predicting the language and dialect of the utterance among 8 languages across 33 dialects. We participated in Track 1 and Track 2, which restricts the use of additional data and develop from-the-scratch multilingual systems. We presented a novel training approach using Multi-Decoder architecture with phonemic Common Label Set (CLS) as intermediate representation. It improved the performance over the baseline (in the CLS space). We also discuss various methods used to retain the gain obtained in the phonemic space while converting them back to the corresponding grapheme representations. Our systems beat the baseline in 3 languages (Track 2) in terms of WER/CER and achieved the highest language ID and dialect ID accuracy among all participating teams (Track 2).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。