arXiv:2608.17360cs.CRcs.AI2026-08

提出公平评估黑盒越狱攻击的新方法,解决资源不公导致的排名失真问题。

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

论文配图:Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
图 1 · 摘自论文原文
  • 用目标调用次数作为统一评估标准,避免计算资源估算偏差。
  • 11种攻击在不同预算下排名变化大,多数现有方法效率双低。
  • 新攻击ReCode仅需7.19次请求调用,20次目标调用即达85%成功率。

可靠的越狱评估对衡量大模型安全性至关重要,但现有研究仅依赖攻击成功率(ASR),未考虑其对攻击预算的依赖性,导致方法间比较不公平。现有计算感知评估将异构资源简化为浮点运算量(FLOPs),难以估算黑盒模型的资源消耗,且无法捕捉特定资源约束。为此,我们提出公平ASR(Fair-ASR)协议,在共享目标调用预算B下评估黑盒越狱攻击,以目标调用次数为直接可观测、方法无关的对比基准,同时单独追踪攻击者调用次数以分析效率。我们在公平框架下重新评估11种代表性攻击,发现攻击排名随目标调用预算显著变化;在同等目标访问条件下,简单随机扰动和手工模板仍具高度竞争力;且无一种被评估的基于LLM的方法在目标与攻击者调用上均高效。基于此效率差距,我们提出ReCode——一种组合式低预算攻击,结合去敏感重写与由公平评估识别出的两种低成本有效原语。在20次目标调用预算下,ReCode在GPT-5上实现85%的攻击成功率,平均每次请求仅需7.19次攻击者调用,展现出在目标与攻击者调用上的双重高效性。

原文摘要 · Abstract (English)

Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous resources into FLOPs, which is difficult to estimate for black-box models and fails to capture resource-specific constraints. To provide a comparable evaluation basis, we introduce Fair-ASR, an evaluation protocol for black-box jailbreak attacks under shared target-call budgets B, using target calls as a directly observable and method-agnostic comparison axis while tracking attacker calls separately for efficiency analysis. We re-evaluate 11 representative attacks under the Fair-ASR protocol and find that attack rankings change substantially across target-call budgets, simple stochastic perturbations and hand-crafted templates remain highly competitive under equal target access, and no evaluated LLM-driven method is efficient in both target and attacker calls. Motivated by this efficiency gap, we introduce ReCode, a compositional budget-efficient attack that combines desensitization rewriting with two effective low-cost primitives identified by Fair-ASR. Under a budget of 20 target calls, ReCode achieves 85% ASR on GPT-5 while requiring only 7.19 attacker calls per request on average, showing strong efficiency in both target and attacker calls.

大模型安全越狱攻击公平评估效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。