攻击视频世界模型的未来预测,诱导自动驾驶误判路况。
BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving

- 用触发擦除序列污染模型学习的动态规律。
- 仅少量训练数据即让模型误判行人消失且道路变空。
- 攻击可迁移至下游决策模块,适合安全评估研究者关注。
视频世界模型在自动驾驶中用于预测场景演化并生成时空表征以支持后续动作规划。这些表征直接影响车辆路径规划,是关键的安全敏感组件。尽管前景广阔,其训练阶段的安全风险仍鲜有研究。本文提出 BadDreamer,一种针对感知环节的可迁移时空后门攻击。不同于传统攻击图像标签或输出内容,BadDreamer 污染视频世界模型所学的转移动态:构建触发-擦除序列——当前帧可见一辆黄色送货骑手,未来帧中该骑手被擦除。仅需少量此类序列微调后,受损模型会隐式建立条件关联:当真实触发出现时,它将幻觉生成骑手消失、道路清空的未来画面。实验表明,这种被污染的未来表征可传递至下游动作模块,无需修改轨迹标签,即可导致非规避性危险路径预测。我们在一个开源感知-动作流水线中验证该攻击,揭示了视频世界模型在表示层面存在安全风险,强调需在干净生成质量之外引入后门检测机制。
原文摘要 · Abstract (English)
Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representations for downstream action prediction. In perception-to-action pipelines, these representations can directly influence ego-vehicle waypoint planning, making the learned future dynamics a critical security-sensitive component. Despite their promise, the training-time security risks of autonomous-driving video world models remain largely unexplored. We present BadDreamer, a transferable spatio-temporal backdoor attack that targets the perception side of this pipeline. Unlike conventional backdoors that manipulate image labels, prompt outputs, or action supervision, BadDreamer poisons the learned transition dynamics of a video world model. It constructs trigger-erasure sequences in which an oncoming yellow delivery rider is visible in the observed context frames but erased from the future frames. After fine-tuning on a small fraction of such sequences, the compromised world model learns a hidden conditional association: when the physical trigger appears, it hallucinates a future where the rider disappears and the road appears clear. We further show that this corrupted future-aware representation can transfer to the downstream action module without directly modifying ego-trajectory labels, inducing unsafe non-evasive waypoint predictions. Our experiments instantiate this attack on a representative open-source perception-to-action pipeline, revealing a representation-level safety risk in autonomous-driving video world models and highlighting the need for backdoor-aware validation beyond clean generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。