首个面向具身机器人对抗攻击的标准化评估基准,兼顾安全与指令执行能力。
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents

- 构建基于国际标准的18类安全违规分类体系,覆盖具身AI风险场景。
- 设计意图对比数据管道,同步评估攻击成功率与正常指令响应能力。
- 提供统一评测流程和开源榜单,支持新攻防方法的持续集成与比较。
视觉语言模型(VLM)的进步催生了新型具身人工智能系统,将模型部署于机器人等物理平台,实现对视觉场景的理解与自然语言指令的执行。现有研究虽提出针对具身AI的越狱攻击与防御机制,但其评估依赖非标准数据集、有限指标,重攻击成功率而忽视安全与功能之间的权衡。现有基准或针对传统对话模型,或仅关注非对抗性安全,无法全面反映具身AI中越狱攻击的输入、后果与评估需求。本文提出RoboJailBench,包含三大核心组件:首先,基于ISO标准、监管规则及真实事件建立18类具身AI安全违规后果分类;其次,设计意图对比数据流水线,通过成对的对抗性与良性目标增强现有数据集,以同时衡量安全性与实用性;最后,构建可演进的资源库,提供标准化指标与统一评估流程,用于整合新攻击与防御方法。基于该基准,我们构建了平衡的分类数据集,并扩展了五个现有数据集;集成四种攻击与两种防御,评估其在主流具身VLM上的表现。本工作首次为具身AI越狱攻击提供标准化评估框架,推动未来研究发展。代码、数据与排行榜已公开,详见 https://purseclab.github.io/benchmark-for-robotics-security。
原文摘要 · Abstract (English)
Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical platforms, e.g. robots and autonomous vehicles, to interpret visual scenes and execute natural language commands in diverse environments. Previous research has introduced jailbreak attacks and defenses for embodied AI. Their evaluations, however, rely on ad-hoc datasets, limited metrics, and emphasize attack success while neglecting the trade-off between security and the ability to follow benign commands. Existing benchmarks and evaluation frameworks either target traditional chat-based models or focus on non-adversarial safety evaluation for embodied AI; neither captures the adversarial risks, inputs, consequences, and evaluation criteria necessary for jailbreak attacks in embodied AI systems. In this paper, we address this gap with RoboJailBench, which consists of three core components. We establish a security taxonomy derived from ISO standards, regulatory rules, and documented incidents. This effort yields 18 categories of security violation consequences for embodied AI. We introduce an intent contrast dataset pipeline that augments existing datasets with paired adversarial and benign goals to measure both security and utility. Lastly, we provide an evolving repository with standardized metrics and a unified process for assessing and integrating new attacks and defenses. With this benchmark, we construct a new taxonomy-balanced dataset and augment five existing datasets. We integrate four attacks and two defenses to evaluate their performance on leading embodied VLMs. This benchmark provides the first standardized evaluation framework for jailbreak attacks in embodied AI and supports future research. We release our code, datasets, and artifacts, and maintain a leaderboard at https://purseclab.github.io/benchmark-for-robotics-security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。