arXiv:2509.07195eess.AS2025-09中稿 · ASRU2025被引 2

解决语音识别在噪声下误判还自信的问题,提升可靠性。

Identifying and Calibrating Overconfidence in Noisy Speech Recognition

  • 后处理校准框架,仅对高自信错误项调整置信度
  • 低信噪比下错误率降低58%,置信度更可信
  • 无需修改模型,适合部署在现有ASR系统

现代端到端语音识别模型如Whisper在噪声环境下不仅识别准确率下降,还表现出过度自信——对错误预测赋予过高置信度。我们系统分析了Whisper在加性噪声下的表现,发现低信噪比时过度自信的错误显著增加,约10%-20%的词元被错误预测且置信度高于0.7。为此,我们提出一种轻量级、后处理的校准框架,可检测潜在过度自信并选择性地对这些词元应用温度缩放,不改变原始模型。在R-SPIN数据集上的评估表明,在低信噪比范围(-18至-5 dB)内,该方法使期望校准误差(ECE)降低58%,归一化交叉熵(NCE)提升三倍,显著改善了极端噪声条件下的置信度可靠性。

原文摘要 · Abstract (English)

Modern end-to-end automatic speech recognition (ASR) models like Whisper not only suffer from reduced recognition accuracy in noise, but also exhibit overconfidence - assigning high confidence to wrong predictions. We conduct a systematic analysis of Whisper's behavior in additive noise conditions and find that overconfident errors increase dramatically at low signal-to-noise ratios, with 10-20% of tokens incorrectly predicted with confidence above 0.7. To mitigate this, we propose a lightweight, post-hoc calibration framework that detects potential overconfidence and applies temperature scaling selectively to those tokens, without altering the underlying ASR model. Evaluations on the R-SPIN dataset demonstrate that, in the low signal-to-noise ratio range (-18 to -5 dB), our method reduces the expected calibration error (ECE) by 58% and triples the normalized cross entropy (NCE), yielding more reliable confidence estimates under severe noise conditions.

语音识别置信度校准噪声鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。