arXiv:2604.05498cs.RO2026-04被引 4

首个评估机器人世界动作模型安全风险的框架,发现其易被恶意指令诱导产生危险行为。

JailWAM: Jailbreaking World Action Models in Robot Control

  • 将不同动作输出统一转为视觉轨迹,实现跨模型风险评估
  • 设计分级安全判别器,可区分安全、动作失败和灾难性风险
  • 结合模拟与快速筛选,降低高成本验证开销,适合安全研究者使用

世界动作模型(WAMs)在机器人操作中展现出强大潜力,能跨任务与环境实现物理交互。但其直接执行高层指令的能力也带来安全隐患,恶意指令可能引发危险行为。为此,我们提出 JailWAM,首个针对 WAM 的越狱攻击评估框架。该框架包含三项创新:首先,引入视觉轨迹映射(Visual-Trajectory Mapping),将模型特异的动作输出转化为统一的视觉轨迹表示,实现对不同架构的持续风险评估;其次,构建由三类安全等级(安全合规、运动失败、灾难性风险)监督的判别器,基于物理后果排序实现细粒度风险识别;第三,设计双路径验证策略,结合快速筛选与闭环物理仿真,仅对潜在高风险样本进行昂贵的实机验证。在 RoboTwin 模拟环境中,JailWAM 对 LingBot-VA 实现了 84.2% 的攻击成功率,表明 WAM 可能易受越狱攻击诱导产生不安全行为。这一发现或推动未来具身系统安全评估与对齐研究。

原文摘要 · Abstract (English)

World Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation, enabling physical interaction across diverse tasks and environments. However, their ability to directly follow high-level instructions and execute physical actions also creates potential safety risks, as adversarially designed instructions may induce unsafe robot behaviors. To systematically assess these risks, we propose JailWAM, the first jailbreak evaluation framework for WAMs. In JailWAM, we integrate three key innovations: Firstly, to address the difficulty of evaluating heterogeneous low-level action outputs, we introduce Visual-Trajectory Mapping, which transforms model-specific actions into unified visual trajectory representations, thereby facilitating consistent risk assessment across WAM architectures. Secondly, to provide efficient and fine-grained assessment of physical risks, we develop a Risk Discriminator supervised by three safety levels ordered according to physical consequence: Safety Compliance, Motion Failure, and Catastrophic Risk. This severity-aware formulation enables the risk discriminator to distinguish different physical outcomes from visual trajectories and support scalable risk screening. Thirdly, to reduce the cost of exhaustively executing adversarial candidates, we design a Dual-Path Verification Strategy that combines rapid risk screening with closed-loop physical simulation, restricting computationally expensive verification to candidates with potential safety risks. Extensive experiments in the RoboTwin simulation environment show that JailWAM achieves an 84.2% attack success rate on LingBot-VA, which indicates that WAMs may be susceptible to jailbreak attacks that induce unsafe physical behaviors. Our findings may motivate further research on the safety evaluation and alignment of future embodied robotic systems.

机器人安全越狱攻击世界模型风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。