BlackDAN多目标优化,让越狱攻击更隐蔽、更相关、成功率更高
BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models
- 用多目标进化算法同时优化成功率、隐蔽性和语义相关性
- 在多个LLM上实现更高成功率且响应更自然不被检测
- 支持用户自定义偏好,适合研究安全与防御的学者
大型语言模型(LLMs)虽具备强大能力,但面临越狱攻击等安全风险,此类攻击利用漏洞绕过防护生成有害内容。现有方法多聚焦于最大化攻击成功率(ASR),常忽略响应与查询的相关性及隐蔽性,导致攻击效果差或易被识别。本文提出BlackDAN,一种基于黑盒的多目标优化越狱框架,采用NSGA-II算法,在多个目标如ASR、隐蔽性与语义相关性间进行优化。通过突变、交叉和帕累托支配机制,实现可解释的越狱提示生成。该框架支持用户根据偏好定制,平衡危害性、相关性等因素。实验表明,BlackDAN在多种LLMs及多模态LLMs上均优于传统单目标方法,兼具更高成功率、更强鲁棒性,且生成响应更具上下文相关性且更难被检测。
原文摘要 · Abstract (English)
While large language models (LLMs) exhibit remarkable capabilities across various tasks, they encounter potential security risks such as jailbreak attacks, which exploit vulnerabilities to bypass security measures and generate harmful outputs. Existing jailbreak strategies mainly focus on maximizing attack success rate (ASR), frequently neglecting other critical factors, including the relevance of the jailbreak response to the query and the level of stealthiness. This narrow focus on single objectives can result in ineffective attacks that either lack contextual relevance or are easily recognizable. In this work, we introduce BlackDAN, an innovative black-box attack framework with multi-objective optimization, aiming to generate high-quality prompts that effectively facilitate jailbreaking while maintaining contextual relevance and minimizing detectability. BlackDAN leverages Multiobjective Evolutionary Algorithms (MOEAs), specifically the NSGA-II algorithm, to optimize jailbreaks across multiple objectives including ASR, stealthiness, and semantic relevance. By integrating mechanisms like mutation, crossover, and Pareto-dominance, BlackDAN provides a transparent and interpretable process for generating jailbreaks. Furthermore, the framework allows customization based on user preferences, enabling the selection of prompts that balance harmfulness, relevance, and other factors. Experimental results demonstrate that BlackDAN outperforms traditional single-objective methods, yielding higher success rates and improved robustness across various LLMs and multimodal LLMs, while ensuring jailbreak responses are both relevant and less detectable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。