arXiv:2605.17971cs.CRcs.AI2026-05

通过优化混淆采样,高效突破大模型安全防护

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

论文配图:Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
图 1 · 摘自论文原文
  • 基于可解释的攻击模型,系统化采样混淆文本
  • 在40次查询内使GPT-4o攻击成功率升至82.67%
  • 适用于模型安全测试与红队演练

尽管经过严格的安全对齐,大语言模型仍易受越狱攻击。现有黑盒方法多依赖启发式模板或穷举尝试,缺乏机制可解释性与查询效率。本研究揭示了模型安全机制中的内在漏洞:安全对齐依赖少数稀疏分布的注意力头,导致大量表征空间监控薄弱。我们构建数学越狱模型,刻画有效文本混淆的微妙边界,并解析观察到的越狱行为。基于此模型,提出Babel框架,通过迭代反馈驱动的混淆分布优化采样,实现无需模型内部信息的高效黑盒攻击。在前沿商用模型上的全面评估表明,Babel在攻击成功率与查询效率上均达领先水平。相比现有方法,其将GPT-4o的攻击成功率从41.33%提升至82.67%,Claude-3-5-haiku从38.33%提升至78.33%,平均仅需40次查询,为大模型安全研究提供强大红队工具。

原文摘要 · Abstract (English)

Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive trials, lacking mechanistic interpretability and query efficiency. In this study, we investigate an intrinsic vulnerability in the safety mechanisms of LLMs, where safety alignment relies on a small set of sparsely distributed attention heads, leaving much of the representational space weakly monitored. We formalize this phenomenon with a mathematical jailbreaking model that characterizes the delicate boundary of effective text obfuscation and analytically explains observed jailbreak behaviors. Guided by this model, we propose Babel, an efficient black-box attack framework that exploits the identified safety gap through systematic obfuscation sampling with iterative, feedback-driven distribution refinement, enabling reliable and high-success jailbreak attacks without access to model internals. Comprehensive evaluations on frontier commercial models demonstrate that Babel achieves state-of-the-art attack success rates and superior query efficiency. Specifically, compared to state-of-the-art methods, Babel increases the attack success rate on GPT-4o from 41.33% to 82.67% and on Claude-3-5-haiku from 38.33% to 78.33% within an average of 40 queries, providing a robust red-teaming methodology for LLMs safety research.

越狱攻击安全对齐黑盒攻击红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。