用偏好优化自动生成隐蔽攻击提示,突破对齐大模型的安全防线。
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
- 通过偏好优化训练攻击模型,自动生成隐蔽的越狱提示。
- 在多个基准上实现更高成功率,且比基线更高效、通用和抗防御。
- 揭示复杂模板更强攻击力,隐晦问题转换更易绕过安全检测。
对齐人类反馈的大语言模型虽受关注,但仍易受越狱攻击,攻击者可通过操控提示诱导有害输出。现有方法多依赖手工模板或生成优化,存在可扩展性差、效率低和通用性不足的问题。为此,我们提出JailPO,一种新颖的黑盒越狱框架,用于检验对齐大模型的安全性。为提升可扩展性和通用性,JailPO精心训练攻击模型以自动产生隐蔽的越狱提示,并引入基于偏好优化的攻击方法,显著提高攻击有效性。我们设计了三种灵活的越狱模式,实验表明,JailPO在自动化攻击的同时保持高效果,且在效率、通用性和防御鲁棒性方面优于基线。分析显示,复杂模板攻击强度更高,而隐蔽问题转换更易引发危险响应并更可能绕过防御机制。
原文摘要 · Abstract (English)
Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipulate prompts to induce harmful outputs. Exploring jailbreak attacks enables us to investigate the vulnerabilities of LLMs and further guides us in enhancing their security. Unfortunately, existing techniques mainly rely on handcrafted templates or generated-based optimization, posing challenges in scalability, efficiency and universality. To address these issues, we present JailPO, a novel black-box jailbreak framework to examine LLM alignment. For scalability and universality, JailPO meticulously trains attack models to automatically generate covert jailbreak prompts. Furthermore, we introduce a preference optimization-based attack method to enhance the jailbreak effectiveness, thereby improving efficiency. To analyze model vulnerabilities, we provide three flexible jailbreak patterns. Extensive experiments demonstrate that JailPO not only automates the attack process while maintaining effectiveness but also exhibits superior performance in efficiency, universality, and robustness against defenses compared to baselines. Additionally, our analysis of the three JailPO patterns reveals that attacks based on complex templates exhibit higher attack strength, whereas covert question transformations elicit riskier responses and are more likely to bypass defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。