CausalDrive让自动驾驶模拟器实时预测因果互动,支持文本控制驾驶社会行为。
CausalDrive: Real-time Causal World Models for Autonomous Driving

- 基于初始画面和车辆轨迹,通过文本指令控制复杂交互的因果生成
- 实现12帧/秒的实时渲染,显著降低扩散模型延迟
- 适合用于强化学习训练、闭环评估与人机协同仿真
世界模型已成为提升自动驾驶数据规模的有前景范式,但现有视频生成模型难以作为交互式模拟器使用。布局条件渲染器依赖所有背景智能体的未来轨迹(即‘预言’),导致其完全非响应;而纯动作条件预测器缺乏语义控制能力,且扩散延迟过高,阻碍闭环策略学习。为弥合这一差距,我们提出CausalDrive,一种可控制、实时运行的基础驾驶世界渲染器。该模型仅需初始前视图、自车轨迹及宏观文本提示,不依赖未来非自车(NPC)布局,迫使模型内生预测因果互动,实现对驾驶社会学的文本驱动控制,可动态编排相同自车动作下的多种反事实反应。为突破效率瓶颈并缓解自回归生成中的协变量偏移问题,我们提出新型上下文强制型DMD架构,结合连续流匹配与自校正蒸馏目标,达到12 FPS的交互速度。这一突破使被动视频生成器转变为可玩神经模拟器。我们在三个下游任务中验证其通用性:(1) 生成式闭环评估,碰撞伪影显著减少;(2) 基于Video2Reward模块的大规模强化学习后训练;(3) 实时人机协同仿真。大量实验表明,在CausalDrive反应式场景中训练的策略在真实世界中展现出更优交互能力。
原文摘要 · Abstract (English)
World models have emerged as a promising paradigm for scaling autonomous driving (AD) data, yet existing video generative models fall short as interactive simulators. Layout-conditioned renderers rely on "oracle" future trajectories of all background agents, rendering them strictly non-reactive. Conversely, pure action-conditioned predictors lack semantic control over complex interactions and suffer from prohibitive diffusion latencies, hindering closed-loop policy learning. To bridge this gap, we present CausalDrive, a controllable, real-time foundation driving world renderer. CausalDrive operates solely on the initial front-view frame, the ego-vehicle's trajectory, and a macroscopic text prompt. By excluding future NPC layouts, we compel the model to intrinsically predict causal interactions, enabling text-driven control over Driving Sociology, allowing users to dynamically orchestrate diverse counterfactual reactions to identical ego-actions. To overcome the efficiency bottleneck and address the covariate shift in autoregressive generation, we propose a novel Context-Forced DMD architecture. This combines continuous flow-matching with a self-correcting distillation objective, achieving interactive speeds of 12 FPS. This breakthrough transforms the passive video generator into a playable neural simulator. We demonstrate its versatility across three downstream applications: (1) generative closed-loop evaluation with significantly mitigated collision artifacts, (2) large-scale Reinforcement Learning (RL) post-training driven by a Video2Reward module, and (3) real-time human-in-the-loop simulation. Extensive experiments validate that policies trained within CausalDrive's reactive scenarios exhibit superior interaction capabilities in the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。