研究印度方言地区数据对语音识别性能的影响
A study on the impact of region specific data on the performance of Indic ASR

- 用单区数据训练,跨区域测试验证泛化能力
- 地理距离越远,错误率越高,相关性显著
- 强调需用多地区数据提升模型鲁棒性
自动语音识别(ASR)系统广泛部署于语言多样性地区,但其在细粒度地理差异下的泛化能力仍缺乏深入研究。本文针对印度语言开展跨县语音识别泛化系统性研究,通过微调作为可控探针,以单一县域语音数据训练模型,并在同语言的其他县域上评估性能。分析多个训练-测试县对的趋势并量化性能差异。为评估地理影响,采用两种距离度量方法分析词错误率(WER)与县间距离的相关性。结果表明,地理距离与WER存在一致相关性,凸显了区域泛化的挑战,强调在印度语音识别开发与评估中需使用地理多样化的语音数据。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems are widely deployed across linguistically diverse regions, yet their ability to generalize across fine-grained geographic variation remains underexplored. We present a systematic study of cross-district ASR generalization for Indian languages, analyzing the impact of regional variation on performance. Using finetuning as a controlled probe, we train models on speech from a single district and evaluate them on other districts within the same language. We examine trends across multiple train test district pairs and quantify performance differences. To assess geographic effects, we analyze the correlation between WER and inter district distance using two distance measures. Our results show consistent correlations between geographic distance and WER, highlighting the challenges of regional generalization and the need for geographically diverse speech data in ASR development and evaluation in India.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。