arXiv:2607.15207cs.LGcs.RO2026-07被引 2

世界动作模型可能想得对却做得错,本文揭示其致命漏洞并提出新攻击框架。

BadWAM: When World-Action Models Dream Right but Act Wrong

论文配图:BadWAM: When World-Action Models Dream Right but Act Wrong
图 1 · 摘自论文原文
  • 设计统一框架BadWAM,用微小视觉扰动破坏模型想象与执行的同步
  • 攻击使任务成功率从96.5%暴跌至43.1%,且可保持未来预测看似合理
  • 适用于评估机器人控制模型安全性的研究人员和开发者

世界-动作模型(WAMs)作为具身控制的新兴基础:不仅预测动作,还学习将动作生成与未来世界预测耦合的表示。这种耦合常被视为鲁棒性、可解释性和安全性的来源,因为机器人的动作理论上可与其预想的未来对比验证。本文揭示该假设的脆弱性。我们提出BadWAM,一个统一的建模与评估框架,用于研究一种新型的WAM特异性对抗攻击——世界-动作漂移攻击。该攻击通过微小的视觉扰动,破坏模型想象与执行之间的对齐。BadWAM从攻击强度和隐蔽性两个维度刻画攻击面。当攻击者追求破坏时,生成仅影响动作的对抗攻击,直接引导模型执行导致任务失败的动作;当额外追求隐蔽性时,生成保持想象一致的对抗攻击,诱导有害动作偏移但维持未来预测接近原始想象。实验表明,两种攻击均显著降低闭环执行下的任务成功率。例如,动作仅攻击将模型性能从96.5%降至43.1%。想象保持型攻击进一步暴露了WAM的特定弱点:适度的未来保持正则化可在减少想象漂移的同时维持强攻击效果。

原文摘要 · Abstract (English)

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.

世界模型对抗攻击机器人控制安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。