发现大模型评估语音时依赖协议捷径,而非真实听音判断。
Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

- 通过三类评测协议检测模型是否绕过音频直接依赖标签或参考数据
- 多模型在特征描述任务中准确率降至0.10以下,配对比较中固定选同一选项
- 强调需针对每种模型-协议组合设计专用探测方法,避免误判
大型音频语言模型(LALMs)被越来越多地用作语音评估的自动评判者。然而,与人类评分高度一致并不意味着其判断基于真实音频。评判者可能依赖评测协议提供的专业标签或参考数据,绕过听音过程形成捷径。本文审计了三种常见部署协议下的此类协议级捷径:特征蓝图评判(音频被声学特征文本替代)、参考条件评判和成对A/B比较。在六种模型、四个属性上,发现多个LALM存在依赖捷径现象。例如,在特征蓝图评判中,错误的专业标签使五位模型的情感判断准确率降至0.10或以下;在拼接式A/B比较中,Qwen3-Omni-Thinking常固定选择同一位置,无视顺序变化。结果表明,仅凭整体一致性会高估LALM评判的有效性,必须联合评估模型与评测协议,并为每对组合设计匹配的捷径探测器。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data supplied by the evaluation protocol itself, taking a shortcut in place of listening to the audio. In this paper, we audit such protocol-level ``shortcuts'' in LALM judges across three common deployment protocols: feature-blueprint judging, where the audio is replaced by a structured text description of acoustic features, reference-conditioned judging, and pairwise A/B comparison. Across six judges and four attributes, we find that several LALMs rely on protocol-level shortcuts. For example, in feature-blueprint judging, incorrect specialist labels reduce five judges' emotion accuracy to 0.10 or below, and in concatenated A/B comparisons, Qwen3-Omni-Thinking often picks the same slot regardless of order swaps. These results indicate that aggregate agreement can overstate the validity of LALM judges unless the model and the evaluation protocol are assessed jointly, and that each model-protocol pair should be evaluated with a matched shortcut probe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。