通过视频掩码重建提升自动驾驶世界模型的泛化与长时预测能力
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
- 用掩码重建损失结合扩散模型,增强特征层级上下文学习
- 在Nuscenes、OpenDV-2K、Waymo上均超越现有方法,长时预测性能显著
- 适合需要长时序建模与多视角生成的自动驾驶研究者
能够从动作中预测环境变化的世界模型对具备强泛化能力的自动驾驶系统至关重要。当前主流驾驶世界模型主要基于视频预测框架,尽管采用先进扩散生成器可生成高保真视频序列,但受限于预测时长和整体泛化能力。本文提出一种新思路:将生成损失与MAE风格的特征级上下文学习相结合。具体设计包括:(1) 采用更可扩展的扩散变换器(DiT)结构,并引入额外掩码构建任务;(2) 设计与扩散过程相关的掩码标记,处理掩码重建与生成扩散之间的模糊关系;(3) 将掩码构建任务拓展至时空域,使用行级掩码替代MAE中的掩码自注意力机制,并引入行级跨视图模块以适配该设计。基于上述改进,提出MaskGWM:一种基于视频掩码重建的通用驾驶世界模型。包含两个变体:MaskGWM-long(侧重长时序预测)与MaskGWM-mview(专注多视角生成)。在标准基准上的全面实验验证了方法有效性,涵盖Nuscenes正常验证、OpenDV-2K长时滚动测试及Waymo零样本验证。定量指标显示,该方法在多个数据集上显著优于现有驾驶世界模型。
原文摘要 · Abstract (English)
World models that forecast environmental changes from actions are vital for autonomous driving models with strong generalization. The prevailing driving world model mainly build on video prediction model. Although these models can produce high-fidelity video sequences with advanced diffusion-based generator, they are constrained by their predictive duration and overall generalization capabilities. In this paper, we explore to solve this problem by combining generation loss with MAE-style feature-level context learning. In particular, we instantiate this target with three key design: (1) A more scalable Diffusion Transformer (DiT) structure trained with extra mask construction task. (2) we devise diffusion-related mask tokens to deal with the fuzzy relations between mask reconstruction and generative diffusion process. (3) we extend mask construction task to spatial-temporal domain by utilizing row-wise mask for shifted self-attention rather than masked self-attention in MAE. Then, we adopt a row-wise cross-view module to align with this mask design. Based on above improvement, we propose MaskGWM: a Generalizable driving World Model embodied with Video Mask reconstruction. Our model contains two variants: MaskGWM-long, focusing on long-horizon prediction, and MaskGWM-mview, dedicated to multi-view generation. Comprehensive experiments on standard benchmarks validate the effectiveness of the proposed method, which contain normal validation of Nuscene dataset, long-horizon rollout of OpenDV-2K dataset and zero-shot validation of Waymo dataset. Quantitative metrics on these datasets show our method notably improving state-of-the-art driving world model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。