arXiv:2504.08104cs.CRcs.AI2025-04被引 5

用遗传算法优化场景变换,让大模型更隐蔽地生成有害内容

Geneshift: Impact of different scenario shift on Jailbreaking LLM

  • 通过遗传算法自动演化场景变换策略
  • 直接提示失败时,成功率从0%提升至60%
  • 在保持表面安全的前提下诱出详细有害响应

越狱攻击旨在迫使大语言模型执行不受限制的行为,已成为人工智能安全领域的重要挑战。尽管基于词典的评估显示现有方法具备较高的攻击成功率,但这些方法难以生成满足有害请求的详细内容,导致在GPT-based评估中表现不佳。为此,我们提出一种黑盒越狱攻击方法GeneShift,利用遗传算法优化场景变换。首先观察到恶意查询在不同场景变换下表现最佳,据此设计遗传算法以进化和选择混合场景变换策略。该方法能引导模型输出详尽且可操作的有害响应,同时维持看似无害的表象,增强隐蔽性。大量实验表明GeneShift显著优于现有方法。值得注意的是,当直接提示无法成功时,GeneShift将越狱成功率从0%提升至60%。

原文摘要 · Abstract (English)

Jailbreak attacks, which aim to cause LLMs to perform unrestricted behaviors, have become a critical and challenging direction in AI safety. Despite achieving the promising attack success rate using dictionary-based evaluation, existing jailbreak attack methods fail to output detailed contents to satisfy the harmful request, leading to poor performance on GPT-based evaluation. To this end, we propose a black-box jailbreak attack termed GeneShift, by using a genetic algorithm to optimize the scenario shifts. Firstly, we observe that the malicious queries perform optimally under different scenario shifts. Based on it, we develop a genetic algorithm to evolve and select the hybrid of scenario shifts. It guides our method to elicit detailed and actionable harmful responses while keeping the seemingly benign facade, improving stealthiness. Extensive experiments demonstrate the superiority of GeneShift. Notably, GeneShift increases the jailbreak success rate from 0% to 60% when direct prompting alone would fail.

越狱攻击遗传算法LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。