arXiv:2607.28685cs.AIcs.IR2026-07

安全基准测试结果不可直接比较,需明确测试对象与指标。

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

论文配图:Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
图 1 · 摘自论文原文
  • 用官方实现和评分器在22个模型上验证4个安全基准的可靠性。
  • 安全分数与能力正相关,但与对齐安全性负相关,存在矛盾。
  • 多数基准测试结果不一致,需说明具体测试条件和目标行为。

代理安全基准衡量不同行为,其得分常被混用为代理安全性的指标。本文将R-Judge、InjecAgent、AgentHarm、AgentDojo四个基准视为待验证测量工具,在最多22个模型上使用官方实现和作者提供的评分器进行测试,并以统一协议评估MMLU和GPQA作为能力综合指标。首先,任何基于二元轨迹判断的基准在$F_1$评分下,'始终肯定'策略可达到$F_1 = 2π/(1+π)$,在R-Judge中为0.690,高于21个实际具有区分能力的模型。随后,三个广覆盖基准对同一18个模型排名差异显著,这种分歧源于小样本偏差:R-Judge特异性与AgentHarm安全性在$n=7$时相关系数为-0.64,在$n=18$时为+0.02,且四分之一随机抽取的$ n=7 $子集相关系数绝对值≥0.5,接近零。保留有效性取决于选择何种结果。能力与任务成功正相关(ρ=+0.60),但与不对齐安全性负相关(ρ=-0.44,n=21)。在配对的$ n=20 $面板中,该对比差异为Δ=-1.00(95% CI [-1.48, -0.49],p<0.001),并通过留一组织分析和组织聚类自助法检验。在扩展至41个模型的面板中,不对齐相关性减弱至-0.16(95% CI [-0.54, +0.22]),而越狱成功率升至+0.34,但均不显著。AgentHarm在控制能力后与三模板越狱安全性的关联最强(ρ=+0.72),但两个工具均测得有害顺从,表明是收敛效度而非普遍安全性。声明基准名称、评分指标、目标行为及模型面板,是安全声明的最低要求。

原文摘要 · Abstract (English)

Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2π/(1+π)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|ρ| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($ρ{=}{+}0.60$) but correlates negatively with misalignment safety ($ρ{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $Δ{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $ρ{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.

安全评估基准测试对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。