arXiv:2506.11079eess.AScs.AI2025-06中稿 · Interspeech 2025被引 4

用提示词提升儿童朗读识别与错字检测效果

Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts

  • 结合语音与文本知识,用提示词优化Whisper和大模型
  • 儿童朗读识别WER降至5.1%,错字检测F1达0.73
  • 适合教育科技、语音评测方向研究者参考

自动朗读评估可为教师提供高效评分支持。然而,相关系统与应用研究仍有限。本文提出一种新型多模态方法,融合音频与文本资源知识。特别地,探索了使用Whisper与指令微调大语言模型(LLM)结合提示词,在儿童语音识别中的表现,以及在下游错字检测任务中的有效性。结果表明,使用提示词的Whisper与提示词驱动的LLM相比无提示基线模型显著提升性能。最优系统在荷兰语儿童朗读语音上达到当前最佳识别效果,词错误率(WER)为5.1%,较基线9.4%显著降低;同时大幅改善错字检测能力,F1分数从0.39提升至0.73。

原文摘要 · Abstract (English)

Automatic reading aloud evaluation can provide valuable support to teachers by enabling more efficient scoring of reading exercises. However, research on reading evaluation systems and applications remains limited. We present a novel multimodal approach that leverages audio and knowledge from text resources. In particular, we explored the potential of using Whisper and instruction-tuned large language models (LLMs) with prompts to improve transcriptions for child speech recognition, as well as their effectiveness in downstream reading mistake detection. Our results demonstrate the effectiveness of prompting Whisper and prompting LLM, compared to the baseline Whisper model without prompting. The best performing system achieved state-of-the-art recognition performance in Dutch child read speech, with a word error rate (WER) of 5.1%, improving the baseline WER of 9.4%. Furthermore, it significantly improved reading mistake detection, increasing the F1 score from 0.39 to 0.73.

语音识别儿童语音提示工程阅读评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。