用普通图片生成攻击指令,让视觉语言模型失守。
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

- 用自然图像自动发现可复用的攻击锚点,闭环优化攻击策略。
- 在未修改的COCO数据集上攻破率高达71.48%,扩展预算下达90%。
- 首次构建超11万条交互式攻击轨迹数据集,适合安全评估与防御研究。
视觉语言模型(VLMs)通过结合视觉感知与文本生成,拓展了安全对齐系统的攻击面。现有多模态越狱攻击主要依赖精心设计的视觉内容、对抗扰动或图像特异性策略,忽视了良性自然图像中可复用的视觉锚点潜力。为此,我们提出MemJack——一种基于记忆增强的多智能体框架,用于自动化VLM红队测试。MemJack在闭环攻击流程中融合视觉锚点发现、视觉-语义伪装、响应评估、反思引导修复与动态重规划。除攻击生成外,还可将公开图像自动转化为攻击锚点,并构建了包含超过11.3万条交互式多模态越狱轨迹的MemJack-Bench数据集,用于安全评估与防御对齐。在完整未修改的COCO val2017图像上的实证评估显示,MemJack对Qwen3-VL-Plus的攻击成功率达71.48%,扩展预算下提升至90%。相比同类基准,在相同自然图像评测设置下,其攻击成功率最高且平均成功轮次更少,验证了自然图像作为可迁移越狱锚点的有效性,并揭示当前安全对齐VLMs存在显著漏洞。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual perception with text generation. Existing multimodal jailbreak attacks primarily rely on crafted visual content, adversarial perturbations, or image-specific attack strategies, leaving the potential of reusable visual anchors in benign natural images largely unexplored. To address the problem, we introduce MemJack, a memory-augmented multi-agent framework for automated VLM red-teaming with natural images. MemJack combines visual-anchor discovery, visual-semantic camouflage, response evaluation, reflection-guided repair, and dynamic replanning within a closed-loop attack pipeline. Beyond attack generation, MemJack automatically transforms public images into attack anchors, enables the construction of MemJack-Bench, a dataset of over 113,000 interactive multimodal jailbreak trajectories for safety evaluation and defensive alignment. Extensive empirical evaluations across full, unmodified COCO val2017 images demonstrate that MemJack achieves a 71.48\% attack success rate (ASR) against Qwen3-VL-Plus, scaling to 90\% under extended budgets. Compared with representative multimodal jailbreak baselines under the same natural-image evaluation setting, MemJack achieves the highest ASR while requiring fewer mean rounds to success, demonstrating superior jailbreak effectiveness on VLMs. These results demonstrate that benign natural images can act as transferable jailbreak anchors and reveal substantial vulnerabilities in current safety-aligned VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。