arXiv:2601.18076cs.LG2026-01NeurIPS被引 14

警告:现有攻击成功率比较可能无效,需重新审视评估标准。

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

  • 从测量理论出发,提出可比性判断标准
  • 发现多数对比违背有效测量原则
  • 适合评估安全测试方法的学者与实践者

我们指出,基于攻击成功率(ASR)对比得出的系统安全性或攻击方法有效性结论,往往缺乏充分证据支持。通过概念、理论和实证分析,我们表明许多结论源于不可比的对比或低效度测量。核心问题在于:何时攻击成功率可以有意义地比较?本文借鉴社会科学测量理论与推断统计学思想,建立判断数值可比性的理论框架,明确ASR可比与不可比的条件。以越狱攻击为例,详细展示典型的‘苹果对橙子’式比较及测量有效性挑战,强调评估时必须确保测量基础一致。

原文摘要 · Abstract (English)

We argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR) comparisons. We show, through conceptual, theoretical, and empirical contributions, that many conclusions are founded on apples-to-oranges comparisons or low-validity measurements. Our arguments are grounded in asking a simple question: When can attack success rates be meaningfully compared? To answer this question, we draw on ideas from social science measurement theory and inferential statistics, which, taken together, provide a conceptual grounding for understanding when numerical values obtained through the quantification of system attributes can be meaningfully compared. Through this lens, we articulate conditions under which ASRs can and cannot be meaningfully compared. Using jailbreaking as a running example, we provide examples and extensive discussion of apples-to-oranges ASR comparisons and measurement validity challenges.

AI安全红队测试评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。