用跨语言无标签数据,让小模型在低资源语言上达到大模型水平。
Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
- 通过持续预训练和分词优化,用3000小时无标签数据提升小模型性能。
- 300M参数模型在波斯语上超越Whisper Large(15亿参数),其他语言表现也优异。
- 适合关注低成本、高效率语音识别的开发者或资源匮乏语言研究者。
低资源语言的自动语音识别仍受限于标注数据稀缺与先进模型所需的计算资源。本文系统研究了针对低资源语言的跨语言连续预训练,以波斯语、阿拉伯语和乌尔都语为主要案例。通过可扩展的无标签数据采集流程构建了3000小时多语言语料库,采用针对性持续预训练结合形态感知分词,训练出一个3亿参数模型,在性能上媲美规模达其5倍的系统。该模型在波斯语上优于Whisper Large v3(15亿参数),在阿拉伯语和乌尔都语上也取得有竞争力的结果,且使用参数更少、标注数据显著减少。研究挑战了‘识别质量随模型规模增长’的主流假设,揭示数据相关性和策略性预训练对低资源场景更为关键。本工作为包容性语音技术提供了一条实用路径,使资源有限的语言无需依赖大规模算力或专有数据集即可实现高效语音识别。
原文摘要 · Abstract (English)
Automatic speech recognition for low-resource languages remains fundamentally constrained by the scarcity of labeled data and computational resources required by state-of-the-art models. We present a systematic investigation into cross-lingual continuous pretraining for low-resource languages, using Perso-Arabic languages (Persian, Arabic, and Urdu) as our primary case study. Our approach demonstrates that strategic utilization of unlabeled speech data can effectively bridge the resource gap without sacrificing recognition accuracy. We construct a 3,000-hour multilingual corpus through a scalable unlabeled data collection pipeline and employ targeted continual pretraining combined with morphologically-aware tokenization to develop a 300M parameter model that achieves performance comparable to systems 5 times larger. Our model outperforms Whisper Large v3 (1.5B parameters) on Persian and achieves competitive results on Arabic and Urdu despite using significantly fewer parameters and substantially less labeled data. These findings challenge the prevailing assumption that ASR quality scales primarily with model size, revealing instead that data relevance and strategic pretraining are more critical factors for low-resource scenarios. This work provides a practical pathway toward inclusive speech technology, enabling effective ASR for underrepresented languages without dependence on massive computational infrastructure or proprietary datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。