arXiv:2608.19936cs.SDcs.AI2026-08

发现主流ASR模型为榜单优化,实际能力被夸大。

Towards Quantifying Benchmark Optimization in ASR Models

论文配图:Towards Quantifying Benchmark Optimization in ASR Models
图 1 · 摘自论文原文
  • 通过三类行为探测,量化模型对榜单参考文本的依赖程度。
  • 顶级开源模型在音频矛盾时仍照搬参考文本,准确率虚高。
  • 仅靠低秩线性调整或末尾加音频即可操控模型行为,适合评估者警惕。

公开基准是衡量自动语音识别(ASR)模型能力的重要指标。然而,由于基准公开,模型可能被优化以适应特定榜单,而无法泛化到真实数据。本文提出一种量化基准优化的方法,重点关注音频不足以确定参考文本的情况。我们识别出三类行为探测:参考文本不一致、掩码数字恢复和拼写切换,揭示模型在音频模糊、被遮蔽或矛盾时仍能复现基准参考文本片段。结果显示,表现最好的开源模型即使面对矛盾音频,仍会原样输出参考文本。通过多种机制探测,我们发现模型会依赖细微声学线索,优先服从榜单优化策略而非忠实还原音频内容。此外,通过低秩线性控制或简单在音频末尾添加内容,可直接操纵模型行为。总体表明,高性能模型表现出受榜单条件影响的行为,导致基准得分虚高,但未反映真正的通用转录能力。

原文摘要 · Abstract (English)

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

ASR基准优化模型评估语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。