arXiv:2503.07608cs.CVcs.RO2025-03被引 110

用强化学习和推理让视觉语言模型更好规划自动驾驶路径

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning

论文配图:AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
图 1 · 摘自论文原文
  • 基于GRPO的强化学习框架,设计四类针对规划任务的奖励函数
  • 两阶段训练策略提升规划性能与训练效率,超越纯微调方法
  • 训练后涌现出多模态规划能力,提升驾驶安全与智能性

OpenAI o1 和 DeepSeek R1 在数学与科学等复杂领域达到甚至超越人类专家水平,其中强化学习(RL)与推理起关键作用。在自动驾驶领域,尽管端到端模型显著提升了规划性能,仍难以应对长尾问题,主要因常识与推理能力不足。已有研究将视觉语言模型(VLMs)引入自动驾驶,但通常仅使用预训练模型并进行简单的监督微调(SFT),未深入探索专用于规划的训练策略与优化。本文提出 AlphaDrive,一种面向自动驾驶的 VLM 强化学习与推理框架。AlphaDrive 引入四种基于 GRPO 的强化学习奖励,采用结合 SFT 与 RL 的两阶段规划推理训练策略。结果表明,相比仅使用 SFT 或无推理的方法,AlphaDrive 显著提升规划性能与训练效率。此外,我们发现经过强化学习训练后,AlphaDrive 出现部分涌现的多模态规划能力,这对提升驾驶安全与效率至关重要。据我们所知,AlphaDrive 是首个将 GRPO 基础的强化学习与规划推理整合至自动驾驶的系统。代码将开源,以推动后续研究。

原文摘要 · Abstract (English)

OpenAI o1 and DeepSeek R1 achieve or even surpass human expert-level performance in complex domains like mathematics and science, with reinforcement learning (RL) and reasoning playing a crucial role. In autonomous driving, recent end-to-end models have greatly improved planning performance but still struggle with long-tailed problems due to limited common sense and reasoning abilities. Some studies integrate vision-language models (VLMs) into autonomous driving, but they typically rely on pre-trained models with simple supervised fine-tuning (SFT) on driving data, without further exploration of training strategies or optimizations specifically tailored for planning. In this paper, we propose AlphaDrive, a RL and reasoning framework for VLMs in autonomous driving. AlphaDrive introduces four GRPO-based RL rewards tailored for planning and employs a two-stage planning reasoning training strategy that combines SFT with RL. As a result, AlphaDrive significantly improves both planning performance and training efficiency compared to using only SFT or without reasoning. Moreover, we are also excited to discover that, following RL training, AlphaDrive exhibits some emergent multimodal planning capabilities, which is critical for improving driving safety and efficiency. To the best of our knowledge, AlphaDrive is the first to integrate GRPO-based RL with planning reasoning into autonomous driving. Code will be released to facilitate future research.

自动驾驶视觉语言模型强化学习规划推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。