arXiv:2606.16605cs.AI2026-06被引 1

构建首个面向世界模型的对抗鲁棒性评测基准,揭示多层级攻击风险。

ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

论文配图:ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control
图 1 · 摘自论文原文
  • 设计五类白盒损失目标,覆盖策略、价值与潜在动态三层次攻击
  • 发现针对价值估计和潜空间的攻击破坏力堪比直接干扰动作
  • 适合关注机器人安全与代理系统可靠性的研究者参考

世界模型因具备学习潜在动力学以支持规划与决策的能力,被广泛应用于机器人与智能体控制系统。随着其在高安全性场景中部署增多,理解其在对抗性条件下的鲁棒性变得至关重要。然而,现有评估缺乏统一基准来测试世界模型代理在策略、价值和潜在动力学层面的对抗威胁。为此,我们提出 ARB4WM,一个统一的预部署鲁棒性与风险评估框架,用于评估世界模型代理在视觉扰动下的表现。ARB4WM 定义了五个白盒损失目标,涵盖上述三个层面,并研究其与单步或多步扰动策略及时间攻击模式(包括全帧、半序列、稀疏帧暴露)结合时的效果。我们在 MetaWorld 和 DeepMind Control Suite 的 20 个任务上,评估了四种 Dreamer 风格代理的表现。结果表明,针对价值估计、潜变量表示以及 RSSM 动态的攻击可造成与直接策略干扰相当的损害;早期或频繁的扰动尤其有害,而输入级防御在自适应攻击下恢复能力有限。这些发现表明,世界模型的安全性评估应覆盖多层次攻击目标与时间暴露协议,而非仅依赖动作空间鲁棒性。源代码见 https://github.com/zaoanguai/ARB4WM。

原文摘要 · Abstract (English)

World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making. As these systems are increasingly deployed in safety-critical settings, understanding their robustness under adversarial conditions has become essential. However, existing evaluations lack a unified benchmark for testing adversarial threats across the policy, value, and latent-dynamics levels of world-model agents. To fill this gap, we present ARB4WM, a unified evaluation framework for pre-deployment robustness and risk assessment of world-model agents under visual perturbations. ARB4WM defines five white-box loss objectives across these three levels and studies their effects when combined with single-step or multi-step perturbation strategies and temporal attack modes, including full-frame, half-sequence, and sparse-frame exposure. Specifically, we evaluate four Dreamer-style agents across 20 tasks from MetaWorld and the DeepMind Control Suite under different loss objectives, perturbation strategies, and temporal attack modes. Results show that attacks targeting value estimation, latent representations, and RSSM dynamics can be as damaging as direct policy disruption, and that early or frequent perturbations are especially harmful, while input-level defenses provide limited recovery under adaptive attacks. These findings suggest that safety, risk, and reliability assessment for world models should cover multiple component-oriented attack objectives and temporal exposure protocols rather than relying solely on action-space robustness. Source code is available at https://github.com/zaoanguai/ARB4WM.

世界模型对抗攻击机器人安全鲁棒性评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。