为农村比哈尔语女性打造语音识别基准,用少量音频生成合成数据提升识别率。
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
- 仅用每名女性25-30秒音频,生成合成语音扩充数据集。
- 合成数据使语音识别错误率降低4.7个点。
- 解决低资源语言女性数据难获取问题,助力数字包容。
数字包容仍面临挑战,尤其在比哈尔语等低资源语言的农村女性群体中。语音访问农业服务、金融交易、政府项目和医疗信息对她们的赋权至关重要,但现有语音识别系统对此类人群测试不足。为此,我们构建了SRUTI基准,包含农村比哈尔语女性说话者数据。在该基准上评估现有语音识别模型表现不佳,源于数据稀缺,而社会文化障碍使得大规模数据采集困难。为此,我们提出仅需约100名农村女性每人提供25-30秒音频,即可生成合成语音。将此合成数据融入现有数据集后,语音识别错误率(WER)下降4.7个百分点,提供了一种可扩展、侵入性极小的解决方案,推动低资源语言语音识别发展与数字包容。
原文摘要 · Abstract (English)
Digital inclusion remains a challenge for marginalized communities, especially rural women in low-resource language regions like Bhojpuri. Voice-based access to agricultural services, financial transactions, government schemes, and healthcare is vital for their empowerment, yet existing ASR systems for this group remain largely untested. To address this gap, we create SRUTI ,a benchmark consisting of rural Bhojpuri women speakers. Evaluation of current ASR models on SRUTI shows poor performance due to data scarcity, which is difficult to overcome due to social and cultural barriers that hinder large-scale data collection. To overcome this, we propose generating synthetic speech using just 25-30 seconds of audio per speaker from approximately 100 rural women. Augmenting existing datasets with this synthetic data achieves an improvement of 4.7 WER, providing a scalable, minimally intrusive solution to enhance ASR and promote digital inclusion in low-resource language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。