arXiv:2510.06461cs.CL2025-10被引 2

用语言学知识优化分词,让濒危语言语音识别更准更快

Linguistically Informed Tokenization Improves ASR for Underresourced Languages

  • 基于语言学知识设计音素级分词,优于传统拼写分词
  • 音素分词使字错误率(WER)和字符错误率(CER)显著降低
  • 适合语言记录员快速转录濒危语言音频

自动语音识别(ASR)是语言学家开展语言记录的重要工具。然而,现代ASR系统依赖数据密集的Transformer架构,难以应用于资源匮乏的语言。本文在濒危的澳大利亚原住民语言Yan-nhangu上微调wav2vec2 ASR模型,比较音素级与拼写级分词策略的效果。结果表明,基于语言学知识的音素分词显著提升性能,相比基线拼写分词方案,明显降低字错误率(WER)和字符错误率(CER)。此外,我们验证了使用人工校正ASR输出的速度远快于从头手写转录,证明了ASR在语言记录流程中的可行性。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) is a crucial tool for linguists aiming to perform a variety of language documentation tasks. However, modern ASR systems use data-hungry transformer architectures, rendering them generally unusable for underresourced languages. We fine-tune a wav2vec2 ASR model on Yan-nhangu, a dormant Indigenous Australian language, comparing the effects of phonemic and orthographic tokenization strategies on performance. In parallel, we explore ASR's viability as a tool in a language documentation pipeline. We find that a linguistically informed phonemic tokenization system substantially improves WER and CER compared to a baseline orthographic tokenization scheme. Finally, we show that hand-correcting the output of an ASR model is much faster than hand-transcribing audio from scratch, demonstrating that ASR can work for underresourced languages.

语音识别濒危语言分词优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。