首个面向物理世界视觉语言模型的统一红队测试基准,可公平比较各类攻击方法。
REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

- 构建共享场景目标生成流水线,实现跨攻击方法的对抗目标对齐
- 文本与字体注入攻击导致最多功能失效,单次攻击已接近迭代攻击效果
- 适用于评估机器人、自动驾驶等安全关键系统中的视觉语言模型漏洞
视觉语言模型(VLMs)正被广泛用于安全敏感的物理世界智能系统中作为感知-推理核心,其感知或推理错误可能导致不安全决策。尽管已有多种红队测试方法,但评估仍分散于不同数据集、指标和威胁模型之间,难以直接比较,且差异来源不明。现有以聊天机器人为核心的红队基准主要聚焦越狱与内容安全评估,未系统涵盖物理场景中的功能性失效及针对物理世界VLM的攻击方法。为此,我们提出REALM,据知是首个面向物理世界VLM的统一红队测试基准。它整合了12种红队方法、3种模型无关防御和13个VLM,在一致的黑盒威胁模型下使用共享数据集与指标。为对齐不同攻击家族的目标,引入代理式目标生成流程,为每个场景构造共享的、场景特定的、物理真实的攻击目标,实现跨方法公平比较。评估显示:文本与字体注入攻击引发最多失败;多模态协同优化产生最强视觉扰动迁移;单次攻击在远低于成本的情况下逼近迭代方法性能;模型规模本身并不带来抗攻击鲁棒性。代码已开源。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used as perception-reasoning backbones for embodied intelligence in safety-critical physical systems, where perception or reasoning errors can lead to unsafe decisions or actions. Although many red-teaming methods have been developed to probe VLM vulnerabilities, their evaluation remains fragmented across datasets, metrics, and threat models, making direct comparison difficult and obscuring whether observed differences arise from stronger attacks, more vulnerable models, or incompatible evaluation settings. Existing chatbot-centric red-teaming benchmarks mainly standardize jailbreak and content-safety evaluation, but they do not systematically capture physically grounded functional failures or cover red-teaming methods that target physical-world VLMs. This raises the key challenge of comparing diverse attack methods under a unified protocol while targeting the same scenario-specific failures. We introduce REALM, to our knowledge the first unified red-teaming benchmark for physical-world VLMs. REALM integrates 12 red-teaming methods, 3 model-agnostic defenses, and 13 VLMs under a practical black-box threat model with shared datasets and metrics. To align adversarial objectives across attack families, REALM introduces an agentic target-generation pipeline that constructs shared, scenario-specific, and physically grounded attack objectives for each scene, enabling fair comparison of diverse red-teaming methods under aligned adversarial goals. Our evaluation shows that text and typographic injection attacks induce the most failures, multimodal co-optimization yields the strongest visual-perturbation transfer, single-pass attacks approach iterative methods at much lower cost, and model scale alone does not confer adversarial robustness. Code is available at https://github.com/UCF-ML-Research/REALM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。