arXiv:2604.04198cs.CVcs.RO2026-04中稿 · ECCV被引 15

DriveVA让自动驾驶模型零样本泛化,同时生成未来画面与动作轨迹。

DriveVA: Video Action Models are Zero-Shot Drivers

论文配图:DriveVA: Video Action Models are Zero-Shot Drivers
图 1 · 摘自论文原文
  • 用统一潜空间联合生成未来视频和驾驶动作序列
  • 在NAVSIM上达90.9分,跨数据集误差降低超78%
  • 适合需要强泛化能力的自动驾驶规划场景

泛化是自动驾驶的核心挑战,真实部署需应对未见场景、传感器域和环境条件。现有基于世界模型的规划方法虽具出色场景理解与多模态预测能力,但跨数据集和传感器配置的泛化性能仍受限,且松耦合规划常导致视觉想象中视频与轨迹不一致。为此,我们提出DriveVA,一种新型自动驾驶世界模型,通过共享潜在生成过程联合解码未来视觉预测与动作序列。DriveVA继承大规模视频生成模型中的运动动力学与物理合理性先验,捕捉连续时空演化与因果交互模式。采用基于DiT的解码器,联合预测未来动作序列(轨迹)与视频,实现规划与场景演进更紧密对齐。此外引入视频续写策略,强化长时滚动一致性。DriveVA在NAVSIM基准上取得90.9的PDM得分。大量实验表明其具备零样本能力与跨域泛化性:相比最先进世界模型规划器,在nuScenes上平均L2误差和碰撞率分别降低78.9%和83.3%,在Bench2Drive(基于CARLA v2)上分别降低52.5%和52.4%。

原文摘要 · Abstract (English)

Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world-model-based planning methods have shown strong capabilities in scene understanding and multi-modal future prediction, yet their generalization across datasets and sensor configurations remains limited. In addition, their loosely coupled planning paradigm often leads to poor video-trajectory consistency during visual imagination. To overcome these limitations, we propose DriveVA, a novel autonomous driving world model that jointly decodes future visual forecasts and action sequences in a shared latent generative process. DriveVA inherits rich priors on motion dynamics and physical plausibility from well-pretrained large-scale video generation models to capture continuous spatiotemporal evolution and causal interaction patterns. To this end, DriveVA employs a DiT-based decoder to jointly predict future action sequences (trajectories) and videos, enabling tighter alignment between planning and scene evolution. We also introduce a video continuation strategy to strengthen long-duration rollout consistency. DriveVA achieves an impressive PDM-based planning performance of 90.9 PDM score on the NAVSIM benchmark. Extensive experiments also demonstrate the zero-shot capability and cross-domain generalization of DriveVA, which reduces average L2 error and collision rate by 78.9% and 83.3% on nuScenes and 52.5% and 52.4% on the Bench2Drive built on CARLA v2 compared with the state-of-the-art world-model-based planner.

自动驾驶世界模型零样本视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。