arXiv:2603.20198cs.CRcs.CV2026-03被引 2

用视觉内容推理实施隐蔽攻击,突破传统防御限制。

Visual Exclusivity Attacks: Automatic Multimodal Red Teaming via Agentic Planning

  • 设计可自主规划的多轮代理框架,实现全局攻击策略生成
  • 在Claude 4.5 Sonnet上达成46.3%攻击成功率,较基线提升2–5倍
  • 提出新数据集VE-Safety,专门评估高风险视觉理解安全问题

当前多模态红队测试将图像视为恶意载荷的载体,通过文字排版或对抗噪声实现攻击,但结构脆弱,一旦载荷暴露即被标准防御拦截。本文提出视觉独占性(Visual Exclusivity, VE)威胁,其危害仅在对技术图示等视觉内容进行推理时才显现。为系统化利用此威胁,提出多模态多轮代理规划(MM-Plan)框架,将越狱从逐轮响应重构为全局策略合成。该框架通过群体相对策略优化(GRPO)训练攻击者规划器,实现无监督自发现有效策略。为严谨评估此类依赖推理的威胁,构建了人工标注的VE-Safety数据集,填补了高风险技术视觉理解评估的空白。实验显示,MM-Plan在Claude 4.5 Sonnet上达到46.3%攻击成功率,在GPT-5上达13.8%,显著优于基线2–5倍,揭示前沿模型仍易受代理式多模态攻击,暴露出当前安全对齐的关键漏洞。警告:本文含潜在有害内容。

原文摘要 · Abstract (English)

Current multimodal red teaming treats images as wrappers for malicious payloads via typography or adversarial noise. These attacks are structurally brittle, as standard defenses neutralize them once the payload is exposed. We introduce Visual Exclusivity (VE), a more resilient Image-as-Basis threat where harm emerges only through reasoning over visual content such as technical schematics. To systematically exploit VE, we propose Multimodal Multi-turn Agentic Planning (MM-Plan), a framework that reframes jailbreaking from turn-by-turn reaction to global plan synthesis. MM-Plan trains an attacker planner to synthesize comprehensive, multi-turn strategies, optimized via Group Relative Policy Optimization (GRPO), enabling self-discovery of effective strategies without human supervision. To rigorously benchmark this reasoning-dependent threat, we introduce VE-Safety, a human-curated dataset filling a critical gap in evaluating high-risk technical visual understanding. MM-Plan achieves 46.3% attack success rate against Claude 4.5 Sonnet and 13.8% against GPT-5, outperforming baselines by 2--5x where existing methods largely fail. These findings reveal that frontier models remain vulnerable to agentic multimodal attacks, exposing a critical gap in current safety alignment. Warning: This paper contains potentially harmful content.

多模态安全红队测试智能体规划视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。