arXiv:2605.09070cs.CRcs.AI2026-05

单一配置成功率误导评估,应报告攻击分布全貌。

Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success

论文配图:Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
图 1 · 摘自论文原文
  • 提出新指标VSM和UC,衡量攻击在参数空间的表现波动与覆盖范围
  • 实测显示最佳配置漏掉大量攻击面,联合覆盖率达93%以上
  • 建议将分布评估作为参数化越狱攻击的最低标准

许多越狱攻击研究仅报告有限参数组合下的成功率,而实际存在大量可调参数。新论文发布时常仅对比单一配置,导致威胁评估和攻击比较严重失真。越狱攻击常涉及系统提示模板、对话轮次、加密分散度、教学样本等多维参数,其自动安全响应率(ASR)在不同设置间差异显著。仅报告最优配置会丢失两个关键信息:该性能在参数空间中的典型性,以及所忽略的攻击面比例。本文提出两种新度量:变体敏感度(VSM)衡量最优结果与平均表现的偏离程度,联合覆盖率(UC)表示所有配置下触发不安全响应的提示占比。在三个开源模型上对两种攻击家族的实验表明:对于PAIR,在Mistral-7B上最佳模板达69% ASR,但联合覆盖率升至88%;在Qwen3-0.6B上分别为75%与93%。对于bijection在Mistral-7B上最佳配置为81% ASR,36种变体联合覆盖全部100个HarmBench-100测试用例。因此,建议采用分布式报告,将VSM与ASR一同公布,并尽可能枚举变体覆盖范围。

原文摘要 · Abstract (English)

Many jailbreak attack research papers report attack success rates for a limited number of parameter settings, even though there are many combinations of parameter settings that could be used. Further, when new jailbreak papers are released, they often benchmark results against single configurations of existing attacks. This position paper argues such practices are fundamentally insufficient for characterising the threat posed by parameterised jailbreak attacks, and comparing attacks. Most jailbreak attacks expose multiple internal parameters, system prompt templates, conversation rounds, cipher dispersion, teaching shots, and ASR varies substantially across these parameters. Reporting only the best-case configuration discards two pieces of information that defenders genuinely need: how typical that performance is across the variant space, and how much of the attack surface is missed by selecting a single variant. We propose two new measures for jailbreak attacks: the Variant Sensitivity Measure (VSM) and Union Coverage (UC). VSM quantifies how far the best reported ASR deviates from the mean ASR across the tested variant space, UC is the total fraction of prompts resulting in unsafe responses across all tested configurations. We empirically demonstrate the importance of these measures using two attack families across three open-source target models. For PAIR, the best template reaches 69% ASR on Mistral-7B and 75% on Qwen3-0.6B, while UC rises to 88% and 93%, respectively. For bijection on Mistral-7B, the best variant reaches 81% ASR, but the 36-variant union covers 100% of HarmBench-100 prompts. We argue that distributional reporting, publishing VSM alongside ASR and enumerating variant coverage as fully as compute allows, should become the new minimum standard for parameterised jailbreak evaluation.

越狱攻击安全评估指标设计分布分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。