arXiv:2604.24278cs.SDcs.AI2026-04中稿 · InterSpeech 2026

提出新指标RAS,让语音识别更懂何时该‘不答’

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

论文配图:RAS: a Reliability Oriented Metric for Automatic Speech Recognition
图 1 · 摘自论文原文
  • 让语音模型学会在不确定时主动放弃输出
  • 新指标平衡准确与拒答,人类偏好校准参数
  • 适合对可靠性要求高的实际应用场景

自动语音识别系统在嘈杂或模糊环境下常给出自信却错误的转录,误导用户和下游应用。传统基于词错误率的评估只关注准确率,忽视转录可靠性。本文提出一种支持拒答的转录框架,使ASR模型能明确回避不确定片段。为评估拒答下的可靠性,提出RAS指标,平衡转录信息量与错误规避能力,其权衡参数由人类偏好校准。通过监督自举与强化学习训练拒答感知的ASR模型,实验表明在保持竞争力准确率的同时,显著提升转录可靠性。

原文摘要 · Abstract (English)

Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream applications. Standard evaluation based on Word Error Rate focuses solely on accuracy and fails to capture transcription reliability. We introduce an abstention-aware transcription framework that enables ASR models to explicitly abstain from uncertain segments. To evaluate reliability under abstention, we propose RAS, a reliability-oriented metric that balances transcription informativeness and error aversion, with its trade-off parameter calibrated by human preference. We then train an abstention-aware ASR model through supervised bootstrapping followed by reinforcement learning. Our experiments demonstrate substantial improvements in transcription reliability while maintaining competitive accuracy.

语音识别可靠性评估拒答机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。