arXiv:2510.15446cs.RO2025-10被引 5

VDRive用视觉语言动作模型加扩散策略,实现更懂环境的端到端自动驾驶。

VDRive: Leveraging Reinforced VLA and Diffusion Policy for End-to-end Autonomous Driving

  • 结合VLA与扩散策略,分上下文和几何两路建模驾驶状态与动作。
  • 在Bench2Drive和nuScenes上达到当前最优性能,闭环与开环均表现优异。
  • 适合研究端到端自动驾驶、可解释决策系统的学者与工程师。

在自动驾驶中,动态环境和极端场景对车辆状态理解与决策鲁棒性构成重大挑战。本文提出VDRive,一种新型端到端自动驾驶框架,通过显式建模状态-动作映射来应对这些挑战,实现可解释且鲁棒的决策。该方法利用视觉语言动作模型(VLA)在状态理解上的进展,结合基于生成式扩散策略的动作头,使驾驶决策兼具上下文与几何合理性。上下文层面,VLA通过令牌生成预训练预测未来观测,观测由条件向量量化变分自编码器(CVQ-VAE)离散编码表示;几何层面,通过强化学习微调VLA,依据当前驾驶条件预测未来轨迹与动作。VLA提供当前状态令牌及预测状态令牌,供动作策略头生成分层动作与轨迹。策略训练中,学习到的评判器评估策略生成动作并提供基于梯度的反馈,形成基于强化学习的策略学习框架。实验表明,VDRive在Bench2Drive闭环基准和nuScenes开环规划任务中均达到当前最优性能。

原文摘要 · Abstract (English)

In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's state understanding and decision making. We introduce VDRive, a novel pipeline for end-to-end autonomous driving that explicitly models state-action mapping to address these challenges, enabling interpretable and robust decision making. By leveraging the advancement of the state understanding of the Vision Language Action Model (VLA) with generative diffusion policy-based action head, our VDRive guides the driving contextually and geometrically. Contextually, VLA predicts future observations through token generation pre-training, where the observations are represented as discrete codes by a Conditional Vector Quantized Variational Autoencoder (CVQ-VAE). Geometrically, we perform reinforcement learning fine-tuning of the VLA to predict future trajectories and actions based on current driving conditions. VLA supplies the current state tokens and predicted state tokens for the action policy head to generate hierarchical actions and trajectories. During policy training, a learned critic evaluates the actions generated by the policy and provides gradient-based feedback, forming an actor-critic framework that enables a reinforcement-based policy learning pipeline. Experiments show that our VDRive achieves state-of-the-art performance in the Bench2Drive closed-loop benchmark and nuScenes open-loop planning.

端到端驾驶视觉语言动作扩散策略强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。