arXiv:2609.04232eess.AScs.CL2026-09

用ASR加速粤语口述史整理,虽有错漏但省时百倍。

Automatic Speech Recognition for Multilingual Oral History Research

  • 用Whisper模型做粤语口述史自动转录,降低人工成本。
  • 最佳配置下词错误率12.10%,但非英语片段易出错。
  • 适合语言复兴项目快速生成初稿,节省99%人工时间。

本文探讨了语音技术在社区主导的语言遗产保护与复兴中的应用。以新西兰粤语复兴为例,口述史是关键策略。自动语音识别(ASR)工具如Whisper显著加速了长期依赖人力的转录流程。然而,现有研究缺乏对代码混杂语境下ASR效果的评估。基于词错误率(WER),最优Whisper配置在非英语片段识别上存在不足,整体达到12.10的WER。尽管如此,该模型仅需原计划1%的时间即可完成初稿转录,仍为社区项目提供高效支持。

原文摘要 · Abstract (English)

This paper offers a unique perspective on how speech technologies are being adopted by community-led heritage language preservation and revitalisation initiatives. As a community-led language maintenance strategy, oral histories play a crucial role in Cantonese language revitalisation in New Zealand. The development of Automatic Speech Recognition (ASR) toolkits, such as Whisper, have expedited what has often been a resource and time-intensive process of transcribing oral history collections. However, there is limited research into the effectiveness of ASR toolkits when applied to code-switched language contexts. Based on Word Error Rate (WER), the best performing Whisper model configuration achieved a WER of 12.10 at the expense of accurately transcribing unsupported non-English segments. However, Whisper remains a useful tool by providing a first-pass transcription using only 1% of the estimated time otherwise needed for manual transcription.

语音识别口述史粤语社区项目

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。