arXiv:2506.04076cs.CLcs.SD2025-06中稿 · the ISCA SLaTE-202…被引 2

精确标注停顿能显著提升口语转录准确率

Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems

  • 用LoRA微调Whisper模型,基于三类停顿标注方案对比
  • 精确标注停顿使词错误率降至5.5%,相对降低11.3%
  • 适合需要精准纠错的外语口语评估场景

语音评估中的逐字转录要求准确捕捉言语不流畅现象,这对后续错误分析和反馈至关重要。然而,许多自动语音识别系统会丢弃或泛化停顿,导致重要声学细节丢失。本文在Speak & Improve 2025语料库上,使用低秩适配(LoRA)微调Whisper模型,未依赖外部音频训练数据。比较三种标注方案:移除停顿(Pure)、通用标签(Rich)、以及由Gemini 2.0 Flash从已有音视频对中推断出的声学精确填充词(Extra)。挑战系统达到6.47% WER(Pure)和5.81% WER(Extra)。赛后实验表明,使用“Extra”方案微调Whisper Large V3 Turbo后,词错误率降至5.5%,相比“Pure”方案(6.2% WER)相对提升11.3%。结果表明,显式且真实的填充停顿标注能显著提升逐字转录准确率。

原文摘要 · Abstract (English)

Verbatim transcription for automatic speaking assessment demands accurate capture of disfluencies, crucial for downstream tasks like error analysis and feedback. However, many ASR systems discard or generalize hesitations, losing important acoustic details. We fine-tune Whisper models on the Speak & Improve 2025 corpus using low-rank adaptation (LoRA), without recourse to external audio training data. We compare three annotation schemes: removing hesitations (Pure), generic tags (Rich), and acoustically precise fillers inferred by Gemini 2.0 Flash from existing audio-transcript pairs (Extra). Our challenge system achieved 6.47% WER (Pure) and 5.81% WER (Extra). Post-challenge experiments reveal that fine-tuning Whisper Large V3 Turbo with the "Extra" scheme yielded a 5.5% WER, an 11.3% relative improvement over the "Pure" scheme (6.2% WER). This demonstrates that explicit, realistic filled-pause labeling significantly enhances ASR accuracy for verbatim L2 speech transcription.

语音识别停顿标注逐字转录Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。