arXiv:2510.02677cs.AIcs.LG2025-10被引 5

自动优化攻击策略,系统评估多模态模型安全漏洞

ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks

  • 构建自适应红队代理,通过多步推理动态生成攻击策略
  • 在多个基准上攻击成功率超基线52.1%,对Claude-4-Sonnet突破90%
  • 生成3万+实例的大型安全数据集,助力模型鲁棒性提升

随着视觉语言模型(VLMs)广泛应用,其多模态接口引入了新的安全风险,使得安全评估变得复杂而关键。现有红队测试要么局限于少数攻击模式,要么高度依赖人工设计,难以实现对新兴真实世界VLM漏洞的可扩展探索。为此,我们提出ARMS——一种自适应红队代理,能够系统性地对VLM进行全面风险评估。给定目标有害行为或风险定义,ARMS通过增强推理的多步编排,自动优化多样化的红队策略,以有效诱导目标VLM产生有害输出。我们提出了11种新颖的多模态攻击策略,涵盖多种对抗模式(如推理劫持、上下文伪装),并通过模型上下文协议(MCP)集成17种红队算法。为平衡攻击多样性与有效性,设计分层记忆与ε-贪婪探索算法。在实例级和策略级基准上的实验表明,ARMS攻击成功率显著优于基线,平均提升52.1%,在Claude-4-Sonnet上超过90%。其生成的红队实例多样性显著更高,揭示出新型潜在漏洞。基于ARMS,我们构建了包含超过3万条实例、覆盖51个风险类别的大规模多模态安全数据集ARMS-Bench,覆盖真实世界威胁与监管风险。使用ARMS-Bench进行安全微调可显著提升VLM鲁棒性,同时保持通用能力,为应对新兴威胁提供切实可行的安全对齐指导。

原文摘要 · Abstract (English)

As vision-language models (VLMs) gain prominence, their multimodal interfaces also introduce new safety vulnerabilities, making the safety evaluation challenging and critical. Existing red-teaming efforts are either restricted to a narrow set of adversarial patterns or depend heavily on manual engineering, lacking scalable exploration of emerging real-world VLM vulnerabilities. To bridge this gap, we propose ARMs, an adaptive red-teaming agent that systematically conducts comprehensive risk assessments for VLMs. Given a target harmful behavior or risk definition, ARMs automatically optimizes diverse red-teaming strategies with reasoning-enhanced multi-step orchestration, to effectively elicit harmful outputs from target VLMs. We propose 11 novel multimodal attack strategies, covering diverse adversarial patterns of VLMs (e.g., reasoning hijacking, contextual cloaking), and integrate 17 red-teaming algorithms into ARMs via model context protocol (MCP). To balance the diversity and effectiveness of the attack, we design a layered memory with an epsilon-greedy attack exploration algorithm. Extensive experiments on instance- and policy-based benchmarks show that ARMs achieves SOTA attack success rates, exceeding baselines by an average of 52.1% and surpassing 90% on Claude-4-Sonnet. We show that the diversity of red-teaming instances generated by ARMs is significantly higher, revealing emerging vulnerabilities in VLMs. Leveraging ARMs, we construct ARMs-Bench, a large-scale multimodal safety dataset comprising over 30K red-teaming instances spanning 51 diverse risk categories, grounded in both real-world multimodal threats and regulatory risks. Safety fine-tuning with ARMs-Bench substantially improves the robustness of VLMs while preserving their general utility, providing actionable guidance to improve multimodal safety alignment against emerging threats.

多模态安全红队测试对抗攻击大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。