用游戏化设计诱骗多模态大模型主动越狱,成功率超90%。
GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models
- 将恶意意图拆解重构为游戏场景,诱导模型主动探索并完成攻击
- 在多个推理模型上实现85.87%至92.13%的越狱成功率
- 专攻具备思维链的模型,突破传统攻击局限,适合安全评估研究者
多模态大语言模型(MLLMs)虽广泛部署,但其安全对齐在对抗输入下仍脆弱。已有研究表明,增加推理步数可破坏安全机制,使模型生成有害内容。然而,多数现有攻击仅提升视觉任务复杂度,未利用模型自身推理激励,导致在具有思维链(Chain-of-Thought)的推理模型上表现不佳。若模型能像人一样思考,能否影响其认知阶段决策以主动完成越狱?为此,我们提出GAMBIT(Gamified Adversarial Multimodal Breakout via Instructional Traps),一种新型多模态越狱框架。该框架分解并重组有害视觉语义,构建游戏化场景,驱动模型作为参与者探索、重构意图并作答以赢得游戏。由此产生的结构化推理链同时增加视觉与文本任务复杂度,使模型聚焦目标达成,降低安全注意力,最终回答重构后的恶意查询。在多个主流推理与非推理型MLLM上的实验表明,GAMBIT取得高攻击成功率(ASR),在Gemini 2.5 Flash上达92.13%,QvQ-MAX上为91.20%,GPT-4o上为85.87%,显著优于基线方法。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous work has shown that increasing inference steps can disrupt safety mechanisms and lead MLLMs to generate attacker-desired harmful content. However, most existing attacks focus on increasing the complexity of the modified visual task itself and do not explicitly leverage the model's own reasoning incentives. This leads to them underperforming on reasoning models (Models with Chain-of-Thoughts) compared to non-reasoning ones (Models without Chain-of-Thoughts). If a model can think like a human, can we influence its cognitive-stage decisions so that it proactively completes a jailbreak? To validate this idea, we propose GAMBI} (Gamified Adversarial Multimodal Breakout via Instructional Traps), a novel multimodal jailbreak framework that decomposes and reassembles harmful visual semantics, then constructs a gamified scene that drives the model to explore, reconstruct intent, and answer as part of winning the game. The resulting structured reasoning chain increases task complexity in both vision and text, positioning the model as a participant whose goal pursuit reduces safety attention and induces it to answer the reconstructed malicious query. Extensive experiments on popular reasoning and non-reasoning MLLMs demonstrate that GAMBIT achieves high Attack Success Rates (ASR), reaching 92.13% on Gemini 2.5 Flash, 91.20% on QvQ-MAX, and 85.87% on GPT-4o, significantly outperforming baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。