研究发现语音大模型难捕捉低资源语言方言特征,提出78小时标注数据集
Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
- 构建78小时孟加拉语方言语音文本数据集Ben-10,用于方言语音识别研究
- 所有深度学习方法在方言识别中表现差,零样本与微调均不理想
- 方言专用训练可缓解性能下降,适合低资源语言语音建模研究者
传统语音识别研究多依赖标准语形式,而方言识别常被视为微调任务。为探究方言差异对自动语音识别(ASR)的影响,我们构建了一个78小时的标注孟加拉语语音转文本(STT)数据集Ben-10。从语言学和数据驱动视角分析表明,语音基础模型在方言识别中表现不佳,无论在零样本还是微调设置下均存在显著性能下降。所有深度学习方法在方言语音数据上均面临挑战,但针对特定方言的模型训练可有效缓解该问题。本数据集亦可作为资源受限条件下ASR算法的域外(OOD)测试资源。项目相关数据与代码已公开。
原文摘要 · Abstract (English)
Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is treated as a fine-tuning task. To investigate the effects of dialectal variations on ASR we develop a 78-hour annotated Bengali Speech-to-Text (STT) corpus named Ben-10. Investigation from linguistic and data-driven perspectives shows that speech foundation models struggle heavily in regional dialect ASR, both in zero-shot and fine-tuned settings. We observe that all deep learning methods struggle to model speech data under dialectal variations but dialect specific model training alleviates the issue. Our dataset also serves as a out of-distribution (OOD) resource for ASR modeling under constrained resources in ASR algorithms. The dataset and code developed for this project are publicly available
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。