构建尼泊尔语语音数据集,用邻近语言迁移提升低资源语音识别
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
- 用邻近的尼泊尔语做跨语言迁移,避免大规模预训练
- 在5.39小时数据上,错误率从52.54%降至17.59%
- 适合濒危语言研究者与低资源语音技术开发者
尼泊尔语(新里语)是加德满都谷地的濒危语言,因缺乏标注语音资源而长期数字边缘化。本文推出Nwāchā Munā——一个5.39小时的手动转录的天城文新里语语音语料库,并建立首个基于保留字形声学建模的基准。我们探究在极低资源自动语音识别(ASR)场景下,是否可借助地理与语言相近的尼泊尔语进行近距跨语言迁移,媲美大规模多语言预训练。微调尼泊尔语Conformer模型,结合数据增强,将字符错误率(CER)从52.54%的零样本基线降低至17.59%,性能接近使用更少参数的多语言Whisper-Small模型。结果表明,从尼泊尔语进行近距迁移是高效替代大型多语言模型的可行方案。我们公开发布数据集与基准,助力新里语社区数字化并推动相关研究。
原文摘要 · Abstract (English)
Nepal Bhasha (Newari), an endangered language of the Kathmandu Valley, remains digitally marginalized due to the severe scarcity of annotated speech resources. In this work, we introduce Nwāchā Munā, a newly curated 5.39-hour manually transcribed Devanagari speech corpus for Nepal Bhasha, and establish the first benchmark using script-preserving acoustic modeling. We investigate whether proximal cross-lingual transfer from a geographically and linguistically adjacent language (Nepali) can rival large-scale multilingual pretraining in an ultra-low-resource Automatic Speech Recognition (ASR) setting. Fine-tuning a Nepali Conformer model reduces the Character Error Rate (CER) from a 52.54% zero-shot baseline to 17.59% with data augmentation, effectively matching the performance of the multilingual Whisper-Small model despite utilizing significantly fewer parameters. Our findings demonstrate that proximal transfer from Nepali language serves as a computationally efficient alternative to massive multilingual models. We openly release the dataset and benchmarks to digitally enable the Newari community and foster further research in Nepal Bhasha.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。