评估语音识别对异常发音的准确率时,需区分真实转录和意图转录。
What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

- 引入真实发音与意图文本双参考标准,避免单一标注误导评估
- 11种模型在双参考下性能差异显著,排名变化明显
- 提醒研究者根据实际应用选择合适转录方式
语音识别系统在异常发音上表现常不佳。一个常被混淆的因素是:在异常语音识别中存在两种有效转录参考——原文转录(包含重复、拖长等真实发音)和意图转录(去除不流畅后的标准文本),其选择取决于上下文与应用场景。现有评估通常将二者合并为单一真实标签,奖励删除不流畅的模型,忽略对原文忠实度的考量。本文以口吃语音为例,使用原文和意图两种参考,对11个来自编码器-解码器、CTC和转换器家族的ASR模型进行基准测试。定量分析揭示了模型在不同转录风格下的性能差异及排名变化。研究强调,模型选择应依据具体应用场景,合理选用转录参考,尤其在异常语音识别中至关重要。
原文摘要 · Abstract (English)
ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcription references: verbatim (actual produced speech, including repetitions/prolongations) and intended (the canonical form of the text with disfluencies removed) in atypical speech recognition depending on context and use-case. Most ASR evaluations conflate this duality into a single ground truth and reward systems that delete disfluencies, ignoring verbatim faithfulness. We benchmark 11 ASR models from encoder-decoder, CTC and transducer families using both verbatim and intended references on atypical stuttered speech as a case study. Our quantitative assessment underlines the disparity in model performance and rankings using the two transcript styles. Through this analysis, we highlight the importance of selecting a suitable transcription reference for valid model selection depending on the use-case, particularly for atypical ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。