arXiv:2503.01865cs.LGcs.AI2025-03ACL被引 16

移除多余约束,让越狱攻击更易跨模型生效

Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints

  • 去掉响应模式和尾部词元约束,提升攻击迁移性
  • 在多个目标模型上成功率从18.4%升至50.3%
  • 增强攻击稳定性和可控性,适合安全研究者参考

越狱攻击能有效诱导大语言模型产生不安全行为,但其在不同模型间的迁移能力有限。本研究聚焦基于梯度的越狱方法,通过分析优化过程,提出新框架揭示迁移性瓶颈。发现响应模式约束与词元尾部约束是导致迁移性低的关键冗余限制。移除这些非必要约束后,显著提升攻击的迁移性与可控性。以Llama-3-8B-Instruct为源模型,在一组具有不同安全等级的目标模型上测试,总体越狱成功率(T-ASR)从18.4%提升至50.3%,同时增强源模型与目标模型上的行为稳定性与可控性。

原文摘要 · Abstract (English)

Jailbreaking attacks can effectively induce unsafe behaviors in Large Language Models (LLMs); however, the transferability of these attacks across different models remains limited. This study aims to understand and enhance the transferability of gradient-based jailbreaking methods, which are among the standard approaches for attacking white-box models. Through a detailed analysis of the optimization process, we introduce a novel conceptual framework to elucidate transferability and identify superfluous constraints-specifically, the response pattern constraint and the token tail constraint-as significant barriers to improved transferability. Removing these unnecessary constraints substantially enhances the transferability and controllability of gradient-based attacks. Evaluated on Llama-3-8B-Instruct as the source model, our method increases the overall Transfer Attack Success Rate (T-ASR) across a set of target models with varying safety levels from 18.4% to 50.3%, while also improving the stability and controllability of jailbreak behaviors on both source and target models.

越狱攻击迁移性LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。