用叙事生成对抗图像,让多模态模型自己暴露恶意指令。
STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
- 用三幕剧结构生成带隐藏恶意的图像序列
- 在Gemini模型上实现93.06%攻击成功率
- 无需改写提示词,适合研究模型安全的学者
统一多模态理解与生成模型(UMMs)在理解和生成任务中表现卓越,但其生成-理解耦合机制存在漏洞。攻击者可利用生成功能构造信息丰富的对抗图像,并通过理解功能在单次交互中吸收恶意内容,称为跨模态生成注入(CMGI)。现有攻击方法多局限于单模态且依赖语义漂移的提示重写,未充分挖掘UMM的独特脆弱性。本文提出STaR-Attack,首个无需语义漂移的多轮越狱攻击框架,通过时空上下文构建与目标查询强相关的恶意事件,采用三幕剧理论生成前情与后情场景,将恶意事件作为隐含高潮隐藏其中。攻击分三步:前两轮利用生成能力生成场景图;第三步引入基于图像的问答游戏,将原始恶意问题混入良性候选中,迫使模型根据叙事上下文选择并回答最相关项。实验表明,STaR-Attack在Gemini-2.0-Flash上达到93.06%攻击成功率,显著优于最强基线FlipAttack,揭示了亟待关注的模型安全缺陷。
原文摘要 · Abstract (English)
Unified Multimodal understanding and generation Models (UMMs) have demonstrated remarkable capabilities in both understanding and generation tasks. However, we identify a vulnerability arising from the generation-understanding coupling in UMMs. The attackers can use the generative function to craft an information-rich adversarial image and then leverage the understanding function to absorb it in a single pass, which we call Cross-Modal Generative Injection (CMGI). Current attack methods on malicious instructions are often limited to a single modality while also relying on prompt rewriting with semantic drift, leaving the unique vulnerabilities of UMMs unexplored. We propose STaR-Attack, the first multi-turn jailbreak attack framework that exploits unique safety weaknesses of UMMs without semantic drift. Specifically, our method defines a malicious event that is strongly correlated with the target query within a spatio-temporal context. Using the three-act narrative theory, STaR-Attack generates the pre-event and the post-event scenes while concealing the malicious event as the hidden climax. When executing the attack strategy, the opening two rounds exploit the UMM's generative ability to produce images for these scenes. Subsequently, an image-based question guessing and answering game is introduced by exploiting the understanding capability. STaR-Attack embeds the original malicious question among benign candidates, forcing the model to select and answer the most relevant one given the narrative context. Extensive experiments show that STaR-Attack consistently surpasses prior approaches, achieving up to 93.06% ASR on Gemini-2.0-Flash and surpasses the strongest prior baseline, FlipAttack. Our work uncovers a critical yet underdeveloped vulnerability and highlights the need for safety alignments in UMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。