让视觉语言模型真正理解视觉信息,实现推理与规划的统一。
Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
- 用分步推理引导模型从视觉输入生成规划决策
- 在nuScenes和DriveLM-nuScenes上超越现有最先进方法
- 适合研究自动驾驶中推理与规划融合的学者
面向自动驾驶的大规模视觉语言模型(VLM)正从感知认知向运动规划演进。然而我们发现两个关键挑战:(1) 模型过度依赖历史输入信息,仅通过捷径获得看似优秀的规划结果,未真正理解视觉输入;(2) 思维链(COT)推理过程与运动规划结果严重错位,如何有效利用复杂推理能力提升规划仍缺乏探索。本文以小规模领域特定VLM为基础,提出Drive-R1,旨在打通场景推理与运动规划。首先,在包含长短思维链数据的精心构建数据集上进行监督微调,使模型从视觉输入逐步推理至最终规划决策。随后,在强化学习框架下训练,以预测轨迹和元动作为奖励信号,激励模型发现对规划更有信息量的推理路径。在nuScenes和DriveLM-nuScenes基准上的实验表明,Drive-R1性能显著优于现有最先进VLM。我们认为Drive-R1为自动驾驶中推理与规划的融合提供了有前景的方向,为未来研究与应用提供方法启示。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challenges in this direction: (1) VLMs tend to learn shortcuts by relying heavily on history input information, achieving seemingly strong planning results without genuinely understanding the visual inputs; and (2) the chain-ofthought (COT) reasoning processes are always misaligned with the motion planning outcomes, and how to effectively leverage the complex reasoning capability to enhance planning remains largely underexplored. In this paper, we start from a small-scale domain-specific VLM and propose Drive-R1 designed to bridges the scenario reasoning and motion planning for AD. Drive-R1 first undergoes the supervised finetuning on a elaborate dataset containing both long and short COT data. Drive-R1 is encouraged to reason step-by-step from visual input to final planning decisions. Subsequently, Drive-R1 is trained within a reinforcement learning framework that incentivizes the discovery of reasoning paths that are more informative for planning, guided by rewards based on predicted trajectories and meta actions. Experimental evaluations on the nuScenes and DriveLM-nuScenes benchmarks demonstrate that Drive-R1 achieves superior performance compared to existing state-of-the-art VLMs. We believe that Drive-R1 presents a promising direction for bridging reasoning and planning in AD, offering methodological insights for future research and applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。