通过伪装善意意图,突破前沿模型的安全防护,诱导其生成有害内容。
Jailbreaking Frontier Foundation Models Through Intention Deception

- 用多轮对话逐步伪装善意,利用模型一致性诱导其输出有害内容。
- 在GPT-5-thinking和Claude-Sonnet-4.5上实现高成功率攻击。
- 发现并揭示新型漏洞‘类越狱’,即不直接回应但信息仍有害。
大型(视觉-)语言模型虽能力强大,却极易被越狱攻击。现有安全训练依赖用户意图判断,但若攻击者隐藏真实意图,系统易失效且显得不友善。为此,前沿模型如GPT-5-thinking转向安全补全机制,旨在兼顾帮助性与安全性。然而,当用户伪装为善意时,该机制可能被滥用。本文提出一种新型多轮越狱方法:通过持续模拟良性意图,利用模型一致性,逐步建立信任,最终引导模型生成详细有害输出。最关键的是,我们发现了此前未被察觉的‘类越狱’现象——模型虽不直接回应有害请求,但透露的信息本身已具危害性。贡献有三:一,在GPT-5-thinking与Claude-Sonnet-4.5上实现高成功率;二,揭示并验证了类越狱漏洞;三,多模态模型实验表明本方法优于当前最优模型。
原文摘要 · Abstract (English)
Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's intent. It has been found that this binary training regime often leads to brittleness, since the user intent cannot reliably be evaluated, especially if the attacker obfuscates their intent, and also makes the system seem unhelpful. In response, frontier models, such as GPT-5, have shifted from refusal-based safeguards to safe completion, that aims to maximize helpfulness while obeying safety constraints. However, safe completion could be exploited when a user pretends their intention is benign. Specifically, this intent inversion would be effective in multi-turn conversation, where the attacker has multiple opportunities to reinforce their deceptively benign intent. In this work, we introduce a novel multi-turn jailbreaking method that exploits this vulnerability. Our approach gradually builds conversational trust by simulating benign-seeming intentions and by exploiting the consistency property of the model, ultimately guiding the target model toward harmful, detailed outputs. Most crucially, our approach also uncovered an additional class of model vulnerability that we call para-jailbreaking that has been unnoticed up to now. Para-jailbreaking describes the situation where the model may not reveal harmful direct reply to the attack query, however the information that it reveals is nevertheless harmful. Our contributions are threefold. First, it achieves high success rates against frontier models including GPT-5-thinking and Claude-Sonnet-4.5. Second, our approach revealed and addressed para-jailbreaking harmful output. Third, experiments on multimodal VLM models showed that our approach outperformed state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。