用可验证奖励微调视觉语言模型,让自动驾驶规划更安全可靠。
LaViPlan : Language-Guided Visual Path Planning with RLVR
- 通过可验证奖励强化学习,对齐语言推理与具体驾驶动作。
- 在域内和域外数据上均提升规划性能,尤其在罕见场景表现更好。
- 适合关注自动驾驶安全、多模态决策的工程师和研究者。
自动驾驶中的分布外(OOD)场景带来严峻挑战,因规划器常无法泛化至训练之外的情境,导致不安全或意外行为。视觉语言模型(VLMs)在处理此类场景中展现出潜力,能提供高层次场景理解并实现用户对齐决策。然而,现有VLMs常存在语言推理与低层轨迹生成之间的错位。本文提出LaViPlan框架,利用基于可验证奖励的强化学习(RLVR)对VLM进行微调,以规划为导向的指标优化模型输出。实验表明,LaViPlan在域内及域外数据集上均提升了规划性能。尽管语言一致性在微调后略有下降,但定性评估显示输出仍具连贯性。消融实验分析了采样比例与推理引导的影响,揭示了这些设计选择对性能的关键作用。结果表明,RLVR作为后训练范式,有望实现语言引导推理与行动级规划的对齐。
原文摘要 · Abstract (English)
Out-of-distribution (OOD) scenarios in autonomous driving pose critical challenges, as planners often fail to generalize beyond their training experience, leading to unsafe or unexpected behavior. Vision-Language Models (VLMs) have shown promise in handling such scenarios by providing high-level scene understanding and user-aligned decisions. However, existing VLMs often exhibit a misalignment between their language-based reasoning and the low-level trajectories required for action-level planning. In this paper, we propose LaViPlan, a framework that leverages Reinforcement Learning with Verifiable Rewards (RLVR) to fine-tune VLMs using planning-oriented metrics. Experimental results show that LaViPlan improves planning performance across both in-domain and out-of-domain datasets. While linguistic fidelity slightly decreases after RLVR-based fine-tuning, qualitative evaluation indicates that the outputs remain coherent. We also conduct ablation studies to analyze the effects of sampling ratio and reasoning guidance, highlighting how these design choices influence performance. These findings demonstrate the potential of RLVR as a post-training paradigm for aligning language-guided reasoning with action-level planning in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。