arXiv:2607.21735cs.AIcs.CR2026-07

AI安全评估能证明什么、不能证明什么,有明确边界可计算。

What AI Red-Team Evaluations Can and Cannot Prove

论文配图:What AI Red-Team Evaluations Can and Cannot Prove
图 1 · 摘自论文原文
  • 用可计算的证据上限界定评估有效性边界
  • 小规模基准仅能验证高频危害,对罕见灾难性风险无效
  • 评估价值取决于能否区分假设,而非攻击是否成功

AI红队评估对某些主张成立,对另一些则不成立,其界限是可计算的而非仅凭判断。本文定义评估的证据上限为在固定测试预算下信念变化的最大倍数,推导出基准零结果下的闭式解,并据此精确定位该界限。发现当危害率高于可计算阈值时,适度规模的基准可在既定证据标准下认证某一类别;此时,无失败记录比单一重现失败更具说服力。低于该阈值时,任何可行规模的被动基准均无法在固定评分规则与近似独立试验结构下提供所需的安全证据。两区间的转换具有闭式表达。该界限不局限于基准:以程序的假设条件诱导率表述,涵盖自适应和自动化红队评估,表明决定证据价值的是假设区分能力而非攻击成功率。对八个评估套件的审计显示,当前基准对高频危害足够,但对罕见灾难性危害仍相差数个数量级。安全基准并非无信息量,而是对特定且可计算命题具信息性,关键在于明确定义其所针对的命题。

原文摘要 · Abstract (English)

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.

AI安全评估方法红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。