arXiv:2505.23627cs.LG2025-05被引 4

用提示词让Whisper更准地转录朗读内容并直接检测朗读错误。

Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection

  • 用提示词融合目标文本,提升原样转录准确率
  • 实现端到端朗读错误检测,效果优于现有方法
  • 适合儿童朗读和异常成人语音的错误分析

识别朗读时的错误(即误读)通常采用后处理方式,将自动语音识别(ASR)转录结果与目标文本比对。然而当ASR转录不准时,该方法效果差。为此,我们提出一种新架构,通过提示词引入目标阅读文本,同时优化原样转录和直接错误检测。实验表明:提示词方式在原样转录上优于微调;且可实现端到端误读检测。我们在儿童朗读和成人异常语音两个案例中验证,相比当前最优方法,本方案显著提升了转录准确率与错误检测能力。

原文摘要 · Abstract (English)

Identifying mistakes (i.e., miscues) made while reading aloud is commonly approached post-hoc by comparing automatic speech recognition (ASR) transcriptions to the target reading text. However, post-hoc methods perform poorly when ASR inaccurately transcribes verbatim speech. To improve on current methods for reading error annotation, we propose a novel end-to-end architecture that incorporates the target reading text via prompting and is trained for both improved verbatim transcription and direct miscue detection. Our contributions include: first, demonstrating that incorporating reading text through prompting benefits verbatim transcription performance over fine-tuning, and second, showing that it is feasible to augment speech recognition tasks for end-to-end miscue detection. We conducted two case studies -- children's read-aloud and adult atypical speech -- and found that our proposed strategies improve verbatim transcription and miscue detection compared to current state-of-the-art.

语音识别提示工程错误检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。