arXiv:2501.14788cs.SDcs.CL2025-01被引 1

用众包等方法扩充低资源语言语音数据,提升识别准确率。

Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages

  • 通过众包、伪标签等技术扩展语音数据量
  • 使用小模型实现5.73%(格鲁吉亚)和9.9%(亚美尼亚)的词错误率
  • 付费众包效果最佳,适合资源有限的研究者

本研究探索了利用众包、伪标签、先进数据预处理及开放数据源(如有声书、Common Voice、YouTube)来增加低资源语言语音识别数据量的方法。尽管这些技术在高资源语言中已广泛应用,但在低资源语言中的应用仍不充分。以亚美尼亚语和格鲁吉亚语为例,研究揭示了语言特性和资源条件对方法效果的影响。结果表明,付费众包在成本与质量间取得最佳平衡,优于志愿者众包、开源有声书和未标注数据使用。消融实验显示,基于扩展数据集训练的模型超越现有基线,在相对小型的FastConformer架构下,格鲁吉亚语和亚美尼亚语的词错误率分别达到5.73%和9.9%。研究已公开亚美尼亚语和格鲁吉亚语模型,支持后续研究与实际应用。

原文摘要 · Abstract (English)

This study explores methods to increase data volume for low-resource languages using techniques such as crowdsourcing, pseudo-labeling, advanced data preprocessing and various permissive data sources such as audiobooks, Common Voice, YouTube. While these methods are well-explored for highresource languages, their application for low-resource languages remains underexplored. Using Armenian and Georgian as case studies, we demonstrate how linguistic and resource-specific characteristics influence the success of these methods. This work provides practical guidance for researchers to choose cost-effective and quality-driven dataset extension strategies for low-resource languages. The key takeaway from various data extension approaches is that paid crowd-sourcing offers the best balance between cost and quality, outperforming volunteer crowd-sourcing, open-source audiobooks, and unlabeled data usage. Ablation study shows that models trained on the expanded datasets outperform existing baselines and achieve 5.73% for Gergian and 9.9% for Armenian ASR word error rate using a relatively small FastConformer architecture. We open-sourced both the Armenian and Georgian models to allow further research and practical applications.

语音识别数据增强低资源语言众包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。