arXiv:2608.11629cs.CL2026-08中稿 · Interspeech 2026

让语言学家零代码用ASR自动转录,提升方言记录效率

Easper: An Accessible ASR Pipeline for Language Documentation

论文配图:Easper: An Accessible ASR Pipeline for Language Documentation
图 1 · 摘自论文原文
  • 通过ELAN标注直接云端微调ASR模型,无需编程
  • 优先转录词汇丰富的叙事内容,错误率下降更快
  • 适合缺乏技术背景的语言记录者和小语种研究者

语音转录是语言记录中的关键瓶颈。尽管多语言自动语音识别(ASR)模型如Whisper提供了方案,但田野语言学家常缺乏使用能力。我们提出Easper,一个开源、无代码的工作流,使语言学家可直接通过ELAN标注,利用云资源迭代微调ASR模型。部署ASR还面临冷启动问题:如何选择首批录音以快速建立准确模型。在瓦努阿图三种语言(比斯拉马语、纳夫桑语、努纳语)上评估了转录优先策略。按录音会话微调模型,比较在优先考虑声学清晰度与语言丰富性下的字符错误率(CER)变化轨迹。结果表明,优先转录词汇丰富的叙述并增加音位重复,即使在嘈杂环境下也能更快提升转录质量。

原文摘要 · Abstract (English)

Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present Easper, an open-source, no-code workflow enabling linguists to iteratively fine-tune ASR models via cloud resources directly from ELAN annotations. Deploying ASR also raises a cold start problem: deciding which recordings to transcribe first to bootstrap an accurate model. Using Easper, we evaluate transcription prioritisation strategies on three Vanuatu languages (Bislama, Nafsan, Nguna). We fine-tune models by recording session, comparing Character Error Rate trajectories when prioritising acoustic cleanliness versus linguistic richness. We demonstrate that prioritising lexically rich narratives and increasing acoustic-phonetic repetition, even in noisy environments, leads to faster improvements in transcription quality.

语音识别语言记录无代码小语种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。