arXiv:2510.12796cs.CVcs.AI2025-10被引 111

用世界模型生成密集自监督信号,提升自动驾驶视觉-语言-动作模型的数据利用效率。

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

论文配图:DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
图 1 · 摘自论文原文
  • 通过预测未来图像构建密集自监督信号,缓解动作标签稀疏问题。
  • 在NAVSIM和自建数据集上表现超越基线,数据量越大优势越明显。
  • 适配多种VLA架构,可降低推理延迟,适合实时自动驾驶系统。

在大规模数据上扩展视觉-语言-动作(VLA)模型为实现更泛化的驾驶智能提供了前景。然而,由于模型容量大而监督信号稀疏、维度低,导致其表征能力未被充分利用。为此,我们提出DriveVLA-W0训练范式,利用世界模型预测未来图像,生成密集的自监督信号,促使模型学习驾驶环境的内在动态。该范式可适配两类主流VLA结构:针对离散视觉标记的自回归世界模型,以及针对连续视觉特征的扩散世界模型。基于世界模型学习到的丰富表征,我们引入轻量级动作专家以降低推理延迟,支持实时部署。在NAVSIM v1/v2基准及680倍更大的自研数据集上的大量实验表明,DriveVLA-W0显著优于BEV与VLA基线,并放大了数据缩放律,即随着训练数据规模增大,性能提升速度加快。

原文摘要 · Abstract (English)

Scaling Vision-Language-Action (VLA) models on large-scale data offers a promising path to achieving a more generalized driving intelligence. However, VLA models are limited by a ``supervision deficit'': the vast model capacity is supervised by sparse, low-dimensional actions, leaving much of their representational power underutilized. To remedy this, we propose \textbf{DriveVLA-W0}, a training paradigm that employs world modeling to predict future images. This task generates a dense, self-supervised signal that compels the model to learn the underlying dynamics of the driving environment. We showcase the paradigm's versatility by instantiating it for two dominant VLA archetypes: an autoregressive world model for VLAs that use discrete visual tokens, and a diffusion world model for those operating on continuous visual features. Building on the rich representations learned from world modeling, we introduce a lightweight action expert to address the inference latency for real-time deployment. Extensive experiments on the NAVSIM v1/v2 benchmark and a 680x larger in-house dataset demonstrate that DriveVLA-W0 significantly outperforms BEV and VLA baselines. Crucially, it amplifies the data scaling law, showing that performance gains accelerate as the training dataset size increases.

自动驾驶世界模型数据缩放VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。