提出新攻击框架BadWorld,揭示视觉世界模型的脆弱性。
BadWorld: Adversarial Attacks on World Models

- 自监督速度攻击破坏模型早期去噪过程
- 轨迹自适应双层优化生成通用控制扰动
- 可暴露模型在安全系统中的重大风险
视觉世界模型(VWMs)能从单张上下文图像生成交互式、动作条件的未来序列。然而,这类模型对对抗扰动的鲁棒性仍不明确。传统攻击因缺乏真实未来视频且无法预测用户后续控制而失效。我们提出针对自回归VWMs的标签无关攻击框架BadWorld,系统解决两大限制:首先,设计自监督速度攻击,直接干扰模型早期去噪动态;其次,采用轨迹自适应双层优化,主动挖掘难样本控制序列,生成对控制无关的扰动。在具有连续与离散控制的代表性VWM上评估,BadWorld暴露了严重的结构脆弱性。视觉上难以察觉的对抗图像会引发未来序列的灾难性退化,导致去噪不完全、结构崩溃及控制不一致。这些发现揭示了将VWM部署于安全关键系统的潜在风险,同时提供了一种实用的隐私保护机制。
原文摘要 · Abstract (English)
Visual world models (VWMs) synthesize interactive, action-conditioned rollouts from a single context image. However, it remains an open question how robust these models are to adversarial perturbations. Standard adversarial attacks fail to assess this vulnerability because attackers lack ground-truth future videos and cannot predict subsequent user controls. We introduce BadWorld, a label-free adversarial framework tailored for autoregressive VWMs that systematically overcomes both constraints. First, to bypass the need for future supervision, we propose a self-supervised velocity attack that directly disrupts the early denoising dynamics of the model. Second, to ensure the attack generalizes across unpredictable user actions, we formulate a trajectory-adaptive bi-level optimization that actively mines hard control sequences to forge control-agnostic perturbations. Evaluated on representative VWMs with continuous and discrete controls, BadWorld exposes severe structural fragility. Visually indistinguishable adversarial images reliably trigger catastrophic degradation in future rollouts, leading to incomplete denoising, structural collapse, and control inconsistency. These findings reveal critical risks for deploying VWMs in safety-critical systems while highlighting a practical mechanism for privacy protection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。