arXiv:2607.05133cs.CV2026-07

统一视频与轨迹生成,用视频监督直接提升自动驾驶规划泛化能力。

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

论文配图:UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation
图 1 · 摘自论文原文
  • 单模型融合视频与轨迹生成,通过掩码调控实现跨模态协同。
  • 在NAVSIM上达91.0 PDMS,零样本迁移至nuScenes和Bench2Drive表现优异。
  • 支持视频、轨迹或联合推理,推理速度提升4.3倍且精度不降。

世界动作模型(WAMs)通过未来视频预测作为密集监督,提升了自动驾驶中的动作泛化能力,但尚不清楚何种架构能更好将视频建模优势传递到轨迹生成。现有级联或双迪特(Two-DiT)结构将视频想象与动作预测分离,削弱了视频所学世界动态对轨迹分支的直接影响:动作模型仍可能过拟合数据集特定驾驶先验,而视频模型仅间接正则化规划。我们提出UNIVERSE,一个基于单一掩码调制扩散变压器(Diffusion Transformer)的统一视频-动作模型。通过在共享生成参数中联合训练未来视频潜在表示与自车轨迹标记,使密集视频监督直接塑造轨迹去噪过程,从而增强跨域动作泛化。为确保因果有效性与高效部署,引入模态解耦可见性掩码,共享历史上下文的同时阻断未来视频与轨迹标记间的相互注意力,防止未来目标泄漏,并可在测试时移除未来视频去噪,实现仅轨迹推理,相较联合视频-动作滚动推演提速4.3倍,同时保持相当的规划精度。同一模型亦支持视频仅推理与联合推理。实验表明,UNIVERSE在NAVSIM上取得91.0 PDMS(优于两迪特变体的89.6),并展示出无需微调的强零样本迁移至nuScenes和Bench2Drive的能力,消融实验验证了单迪特统一、视频联合训练及掩码驱动模态解耦的重要性。

原文摘要 · Abstract (English)

World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as dense supervision for scene dynamics and temporal causality. However, it remains unclear which architecture better transfers video-modeling benefits to trajectory generation. Existing cascaded or dual-DiT designs separate video imagination from action prediction, weakening the transfer of video-learned world dynamics to the trajectory branch: the action model may still overfit dataset-specific driving priors, while the video model only indirectly regularizes planning. We propose UNIVERSE, a unified video-action model built upon a single mask-modulated Diffusion Transformer. By co-training future video latents and ego-trajectory tokens within shared generative parameters, UNIVERSE allows dense video supervision to directly shape trajectory denoising, leading to stronger cross-domain action generalization. To ensure causal validity and efficient deployment, we introduce a Modality-Decoupling Visibility Mask, which shares historical context across modalities while blocking mutual attention between future video and trajectory tokens. This prevents future-target leakage and enables trajectory-only inference by removing future-video denoising at test time, achieving a $4.3\times$ speedup over joint video-action rollout while maintaining comparable planning accuracy. The same model also supports video-only and joint video-action rollouts. Experiments show that UNIVERSE achieves 91.0 PDMS on NAVSIM (vs. 89.6 for the Two-DiT variant), and demonstrates strong zero-shot transfer to nuScenes and Bench2Drive without fine-tuning, while ablations confirm the importance of single-DiT unification, video co-training, and mask-based modality decoupling.

自动驾驶视频生成扩散模型轨迹预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。