用少量数据提升濒危语言语音识别,解决田野调查中的转录难题。
Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
- 微调多语言模型MMS与XLS-R,控制训练时长测试效果。
- 数据少于1小时用MMS,超1小时后XLS-R性能相当。
- 为语言学家提供可复现的语音识别适配方法。
自动语音识别(ASR)在高资源语言上已取得优异性能,但在语言田野调查中的应用仍受限。田野录音面临自发语流、环境噪声及极度稀少的标注数据等挑战。本文在五种类型各异的低资源语言上,基准测试了两种微调的多语言ASR模型——MMS与XLS-R,在控制训练数据时长的条件下评估其表现。结果表明:当训练数据极小时,MMS表现更优;当数据量超过1小时,XLS-R达到相当水平。研究提供基于语言学的分析,为语言学家提供可复现的ASR适应策略,助力缓解语言记录中的转录瓶颈。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) has reached impressive accuracy for high-resource languages, yet its utility in linguistic fieldwork remains limited. Recordings collected in fieldwork contexts present unique challenges, including spontaneous speech, environmental noise, and severely constrained datasets from under-documented languages. In this paper, we benchmark the performance of two fine-tuned multilingual ASR models, MMS and XLS-R, on five typologically diverse low-resource languages with control of training data duration. Our findings show that MMS is best suited when extremely small amounts of training data are available, whereas XLS-R shows parity performance once training data exceed one hour. We provide linguistically grounded analysis for further provide insights towards practical guidelines for field linguists, highlighting reproducible ASR adaptation approaches to mitigate the transcription bottleneck in language documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。