攻击视觉-语言-动作模型的未来想象,可绕过安全检测
Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models
- 通过扰动观察值生成虚假未来轨迹,实现对世界模型的隐蔽攻击
- 无目标攻击威力是随机扰动的60倍,检测率达AUC 1.0
- 适合研究模型鲁棒性、安全验证与对抗攻击的学者
近期许多视觉-语言-动作(VLA)策略采用先想象后行动的设计。世界-动作模型(WAM)首先将短期未来以潜在轨迹 z~ 的形式进行想象,后续动作据此生成。我们发现,被信任的想象过程才是攻击的薄弱环节,而非反应式策略。下游的可信判断器(如安全闸门、视觉模型预测控制规划器或想象-验证检查器)会将 z~ 视为未来预测。因此,策略的鲁棒性并不保证依赖 WAM 的系统的安全性。根本原因在于一种不对称性:污染想象容易(仅需将 z~ 移出自然未来流形),但精准操控难(必须到达指定流形上的目标)。本文采用基于能力的威胁模型,假设观测扰动在 L-infinity 范围内。攻击者通过可微的观测到想象映射,使用投影梯度下降进行攻击。同样的非流形特性也启发了一种无需参数的去噪检测器。评估了三个目标模型:RynnVLA-002、LingBot-VA 与 LaDi-WM。无目标污染强度约为随机扰动的60倍,检测 AUC 达 1.0;有目标控制则受限。自适应攻击若要规避检测,就必须放弃污染。反应式策略对被污染的想象仍保持稳健。然而,原生想象驱动的模型预测控制(MPC)首次出现针对性任务失败(ε=0.01 时成功率从 0.70 降至 0.05,Fisher p < 10^-4)。
原文摘要 · Abstract (English)
Many recent vision-language-action (VLA) policies adopt an imagine-then-act design. A world-action model (WAM) first imagines a short future as a latent trajectory z~, on which the action is then conditioned. We identify this trusted imagination, rather than the reactive policy, as the exposed attack surface. A downstream oracle, such as a safety gate, a visual model-predictive-control (MPC) planner, or an imagine-then-check verifier, consumes z~ as a prediction of the future. The robustness of the policy therefore does not entail the robustness of systems that rely on the WAM. The underlying phenomenon is an asymmetry. Corrupting the imagination is easy, since it requires only displacing z~ from its natural-future manifold. Steering it precisely is hard, since it must reach a specified on-manifold target. We adopt a capability-based threat model with an L-infinity-bounded observation perturbation. The attacker applies projected gradient descent through the fully differentiable observation-to-imagination map. The same off-manifold property motivates a parameter-free denoiser detector. We evaluate three targets: RynnVLA-002, LingBot-VA, and LaDi-WM. Untargeted corruption is roughly 60x stronger than random and is detected at AUC 1.0. Targeted control remains bounded. An adaptive attacker evades detection only by forgoing corruption. The reactive policy remains robust to corrupted imagination. A native imagination-driven MPC, however, exhibits the first adversary-specific task failure (at epsilon=0.01, success 0.70 versus 0.05; Fisher p < 10^-4).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。