arXiv:2605.28833cs.CLcs.AI2026-05

提升儿童语音识别准确率,自动筛选可靠转录文本。

Transcribing Children's Speech: ASR Performance and Obtaining Reliable Orthographic Transcriptions

  • 用语音比对法自动筛选正确发音的语句。
  • Whisper模型在儿童语音上误差率最低,达5.54%。
  • 可自动识别42%的可靠转录,大幅减少人工校对。

自动语音识别(ASR)有望显著降低儿童语音研究中的手动标注成本。然而,在低资源语言中,由于缺乏儿童专用预训练模型及噪声条件高度多样,获取高质量的自动转录仍具挑战。本研究通过两个问题评估九种来自Whisper、Parakeet和Wav2Vec2三个模型家族的ASR模型在两个荷兰儿童语音数据集JASMIN和DART上的表现。问题一:分析ASR模型在儿童语音上的效果。微调后的Whisper-medium模型表现最佳,在JASMIN上字错误率(WER)为5.54%,在更嘈杂的DART上为70.37%。问题二:探究能否在无需人工验证的前提下,自动选出可信赖的转录子集。采用基于语句级别的选择方法,将ASR输出与原始朗读提示进行比对,以识别正确发音的录音。使用该方法,可分别自动识别出42.0%(JASMIN)和18.1%(DART)的语句为高置信度正确,其语句级精确率均超过98.3%,显著降低人工校验需求。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) has the potential to substantially reduce manual annotation effort in child speech research by generating automatic transcriptions. However, obtaining reliably high-quality ASR transcriptions for child speech remains challenging in low-resource languages due to limited child-specific pre-trained models and highly diverse noise conditions. This study investigates the effectiveness of state-of-the-art ASR models on child speech through two research questions, by evaluating nine ASR models from three model families (Whisper, Parakeet, and Wav2Vec2) on two Dutch child speech datasets, JASMIN and DART. Research question 1 examines the performance of ASR-models applied to child speech. The fine-tuned Whisper-medium model achieves the best overall performance, with a WER of 5.54% on JASMIN and 70.37% on DART, showing that the noisy DART data are clearly more challenging. Research question 2 examines to what extent it is possible to select a subset for which reliable orthographic transcriptions can be obtained automatically, without the need for manual verification. We use an utterance-level selection method that compares ASR output with the original read prompt to identify correctly pronounced recordings. Using the proposed selection method, 42.0% [for JASMIN] and 18.1% [for DART] of the utterances can be automatically identified as correctly pronounced with high confidence, resulting in very low error rates on an utterance level (precisions of 98.3% and higher) and reducing the need for manual verification.

语音识别儿童语音自动标注转录质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。