arXiv:2511.13548cs.CRcs.AI2025-11被引 1

提出进化框架生成高隐蔽性对抗提示,突破对齐大模型安全防护。

ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models

  • 多层级文本扰动增强攻击多样性,覆盖字符、词、句层面。
  • 基于语义相似度的可解释评估,确保攻击输出既自然又有害。
  • 双维度判断机制降低误报率,适合安全测试与防御研究者使用。

大型语言模型(LLMs)的快速应用带来了变革性应用的同时也引入了新的安全风险,如越狱攻击可绕过对齐保护以诱导有害输出。现有自动化越狱生成方法(如 AutoDAN)存在变异多样性不足、适应度评估浅显、依赖关键词检测易失效等问题。为此,我们提出 ForgeDAN——一种新型进化框架,用于生成语义连贯且高效的对抗性提示,以攻击对齐的 LLM。首先,ForgeDAN 在字符、词和句子层面上引入多策略文本扰动,提升攻击多样性;其次,采用基于文本相似度模型的可解释语义适应度评估,引导进化过程趋向语义相关且有害的输出;最后,集成双维度越狱判定机制,利用基于 LLM 的分类器联合评估模型合规性与输出危害性,从而减少误报并提高检测有效性。实验表明,ForgeDAN 在保持自然性和隐蔽性的前提下,实现高越狱成功率,优于现有最先进方案。

原文摘要 · Abstract (English)

The rapid adoption of large language models (LLMs) has brought both transformative applications and new security risks, including jailbreak attacks that bypass alignment safeguards to elicit harmful outputs. Existing automated jailbreak generation approaches e.g. AutoDAN, suffer from limited mutation diversity, shallow fitness evaluation, and fragile keyword-based detection. To address these limitations, we propose ForgeDAN, a novel evolutionary framework for generating semantically coherent and highly effective adversarial prompts against aligned LLMs. First, ForgeDAN introduces multi-strategy textual perturbations across \textit{character, word, and sentence-level} operations to enhance attack diversity; then we employ interpretable semantic fitness evaluation based on a text similarity model to guide the evolutionary process toward semantically relevant and harmful outputs; finally, ForgeDAN integrates dual-dimensional jailbreak judgment, leveraging an LLM-based classifier to jointly assess model compliance and output harmfulness, thereby reducing false positives and improving detection effectiveness. Our evaluation demonstrates ForgeDAN achieves high jailbreaking success rates while maintaining naturalness and stealth, outperforming existing SOTA solutions.

越狱攻击LLM安全进化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。