arXiv:2607.24054cs.AI2026-07

评估智能体成功原因时,正确答案可能掩盖真实推理路径。

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

  • 通过替换目标值测试成功是否依赖正确信息
  • 实验显示25.9%的得分提升源于正确目标可用性
  • 适合关注评估公平性与模型可解释性的研究者

一个正确的答案可能掩盖智能体成功的真正原因。一旦智能体在评估中改变其信息状态,正确性就无法区分有意推理与答案获取。结果证据和暴露检测无法确定成功是否依赖于获得的目标值,我们称此为缺失的评估对象——成功溯源。AcquaBench 通过在四个标准化任务上进行匹配的 CLEAN、GOLD、SHAM 值替换,并结合联合 qid-聚类分析来审计该问题。CLEAN 保留基准授权信息,GOLD 提供正确目标,SHAM 保持源结构与暴露机会但替换为匹配的错误值。GOLD 减 CLEAN 测量对正确目标可用性的总响应;GOLD 减 SHAM 测试该响应是否超越匹配源暴露而追踪目标正确性。在 D0 中,GOLD 比 SHAM 高出 19.1 至 25.9 个百分点,表明成功依赖于正确值。在 D2 中,即使在分布充分条件下,共处(coloc)也不再是高分标记(AUROC 0.376 vs 0.142),行为依赖仍可能超出探测范围。模型对比中,支持的 5.0 分 CLEAN 分差压缩至原始 GOLD 差值 -0.6 分,未建立排名反转。智能体基准应同时报告成功及其信息状态的支持程度。

原文摘要 · Abstract (English)

A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes intended reasoning from answer acquisition. Outcome evidence and exposure detection do not establish whether success depended on an acquired target; we call this missing evaluation object success provenance. AcquaBench audits it through matched CLEAN, GOLD, and SHAM value substitution on four standardized surfaces with joint qid-clustered analysis. CLEAN retains benchmark-authorized information. GOLD makes the correct target available. SHAM preserves source structure and exposure opportunity but substitutes a matched incorrect value. GOLD minus CLEAN measures the total score response to correct-target availability; GOLD minus SHAM tests whether that response tracks target correctness beyond matched source exposure. In D0, GOLD exceeds SHAM by 19.1 to 25.9 percentage points, showing that success follows the correct value. In D2, GOLD still exceeds SHAM under distributed sufficiency while coloc no longer transfers as a high-score marker, with AUROC 0.376 and 0.142. Behavioral dependence can thus persist beyond this probe's intended observation unit. In model comparison, a supported 5.0-point CLEAN score gap compresses to a raw GOLD difference of -0.6 points without establishing rank inversion. Agent benchmarks should report success together with whether the evaluated information state supported it.

智能体评估成功溯源基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。