arXiv:2501.00055cs.CRcs.AI2025-01被引 9

用病毒进化思路生成高效可迁移的越狱攻击

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models

  • 借鉴病毒进化机制,用大模型自身做演化算子
  • 在多个基准上性能超越或媲美现有攻击方法
  • 适合研究模型安全漏洞与对抗攻击的学者

尽管安全对齐的大语言模型(LLMs)已成为多智能体系统等复杂现实问题求解的核心,但仍易受越狱攻击等对抗性查询影响,诱使生成有害内容。研究攻击方法有助于理解模型局限性,并在有用性和安全性间权衡。然而,现有越狱攻击主要依赖不透明的优化技术(如逐标记梯度下降)和启发式搜索方法(如大模型精炼),在透明性、可迁移性和计算成本方面存在不足。为此,我们受生物病毒进化与感染过程启发,提出基于进化算法的越狱攻击方法 LLM-Virus,即进化越狱。该方法将越狱攻击视为演化与迁移学习问题,利用大模型作为启发式演化算子,实现高攻击效率、强可迁移性与低时间成本。在多个安全基准上的实验结果表明,LLM-Virus 的表现优于或媲美现有攻击方法。

原文摘要 · Abstract (English)

While safety-aligned large language models (LLMs) are increasingly used as the cornerstone for powerful systems such as multi-agent frameworks to solve complex real-world problems, they still suffer from potential adversarial queries, such as jailbreak attacks, which attempt to induce harmful content. Researching attack methods allows us to better understand the limitations of LLM and make trade-offs between helpfulness and safety. However, existing jailbreak attacks are primarily based on opaque optimization techniques (e.g. token-level gradient descent) and heuristic search methods like LLM refinement, which fall short in terms of transparency, transferability, and computational cost. In light of these limitations, we draw inspiration from the evolution and infection processes of biological viruses and propose LLM-Virus, a jailbreak attack method based on evolutionary algorithm, termed evolutionary jailbreak. LLM-Virus treats jailbreak attacks as both an evolutionary and transfer learning problem, utilizing LLMs as heuristic evolutionary operators to ensure high attack efficiency, transferability, and low time cost. Our experimental results on multiple safety benchmarks show that LLM-Virus achieves competitive or even superior performance compared to existing attack methods.

越狱攻击进化算法模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。