研究大模型越狱攻击的计算效率与成功率关系,发现提示类方法最省算力。
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
- 将越狱攻击视为算力受限的优化过程,统一用浮点运算量衡量进展
- 提示类攻击比优化类更高效,且成功度高、隐蔽性好
- 误导类危害最容易诱发,攻击效果显著依赖具体目标类型
大型语言模型仍易受越狱攻击影响,但对其在不同攻击方法、模型家族和危害类型下,攻击成功率如何随投入算力变化的理解仍不系统。本文建立越狱攻击的缩放定律框架,将每次攻击视为算力受限的优化过程,并在统一的浮点运算量(FLOPs)轴上测量进展。系统评估涵盖四类典型越狱范式:基于优化的攻击、自修正提示、采样选择和遗传优化,在多个模型家族与规模下覆盖多样化有害目标。通过拟合简单的饱和指数函数于FLOPs-成功率轨迹,揭示攻击者预算与成功率的关系,并由此得出可比的效率总结。实证显示,提示类方法相比优化类更高效;通过同一状态对比分析,证明提示攻击在提示空间中更有效进行优化。此外,攻击在成功率与隐蔽性上占据不同权衡区域,提示类方法位于高成功、高隐蔽区间。最后发现,模型脆弱性高度依赖目标类型:涉及误导信息的危害通常比其他非误导性危害更容易触发。
原文摘要 · Abstract (English)
Large language models remain vulnerable to jailbreak attacks, yet we still lack a systematic understanding of how jailbreak success scales with attacker effort across methods, model families, and harm types. We initiate a scaling-law framework for jailbreaks by treating each attack as a compute-bounded optimization procedure and measuring progress on a shared FLOPs axis. Our systematic evaluation spans four representative jailbreak paradigms, covering optimization-based attacks, self-refinement prompting, sampling-based selection, and genetic optimization, across multiple model families and scales on a diverse set of harmful goals. We investigate scaling laws that relate attacker budget to attack success score by fitting a simple saturating exponential function to FLOPs--success trajectories, and we derive comparable efficiency summaries from the fitted curves. Empirically, prompting-based paradigms tend to be the most compute-efficient compared to optimization-based methods. To explain this gap, we cast prompt-based updates into an optimization view and show via a same-state comparison that prompt-based attacks more effectively optimize in prompt space. We also show that attacks occupy distinct success--stealthiness operating points with prompting-based methods occupying the high-success, high-stealth region. Finally, we find that vulnerability is strongly goal-dependent: harms involving misinformation are typically easier to elicit than other non-misinformation harms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。