arXiv:2606.15788cs.CRcs.AI2026-06

用遗传算法自动生成绕过LLM安全限制的恶意后缀

GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking

论文配图:GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking
图 1 · 摘自论文原文
  • 基于遗传算法迭代优化对抗性后缀,无需模型内部参数
  • 在黑盒环境下成功突破多个主流LLM的安全防护
  • 揭示现有对齐机制漏洞,适合安全研究者参考

大型语言模型(LLMs)是当前人工智能信息生态系统的核心组件。为防范有害或违反政策的输出,商用系统采用先进的对齐策略和多层内容审核机制。尽管如此,近期研究已证明LLMs仍易受对抗性操纵,尤其是通过越狱和提示注入技术。本文提出GAS-Leak-LLM,一种基于遗传算法的新型越狱攻击方法,通过系统演化对抗性后缀以绕过安全约束。该方法在严格黑盒设置下运行,无需访问模型参数或内部结构,反映实际部署环境中的威胁场景。通过选择、变异和交叉等启发式策略的迭代应用,框架系统探索离散提示空间,识别高适应度的对抗性后缀。实证结果揭示了现有安全执行机制的关键缺陷,并证实了所提攻击的有效性和实用性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) constitute pivotal components within the AI-dominated information technology ecosystem. To mitigate risks associated with harmful or policy-violating outputs, commercial systems employ advanced alignment strategies and multi-layered content moderation mechanisms. Despite these safeguards, recent research has demonstrated that LLMs remain vulnerable to adversarial manipulation, particularly through jailbreaking and prompt injection techniques. In this work, we propose GAS-Leak-LLM a novel jailbreaking attack based on a genetic algorithm that systematically evolves adversarial suffix to bypass safety constraints. Operating in a strict black-box setting, our method requires no access to model parameters or internals, thereby reflecting realistic threat scenarios in deployed systems. Through the iterative application of selection, mutation, and crossover heuristics, the framework systematically explores the discrete prompt space to identify high-fitness adversarial suffixes. Empirical findings reveal critical shortcomings in existing safety enforcement mechanisms and confirm the effectiveness and practical viability of the proposed attack.

越狱攻击遗传算法黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。