arXiv:2512.09349cs.RO2025-12

用视觉语言模型引导强化学习,让自动驾驶更聪明、更可靠。

COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning

  • 用视觉语言模型生成关键交通元素的推理链,提供语义决策先验。
  • 在训练环境中成功率提升30%,未见场景下提升50%。
  • 适合关注自动驾驶可解释性与泛化能力的研究者。

端到端自动驾驶框架在泛化能力、训练效率和可解释性方面仍面临挑战。现有方法虽利用大规模数据上的视觉语言模型(VLM)进行监督学习以增强推理,但在新场景中鲁棒性不足;而基于强化学习(RL)的方法虽适应性强,却数据效率低且决策过程不透明。为此,我们提出COVLM-RL,一种融合关键对象导向(CO)推理与VLM引导强化学习的新型端到端驾驶框架。具体而言,设计了一种思维链(CoT)提示策略,使VLM对关键交通要素进行推理并生成高层语义决策,将多视角视觉输入转化为结构化的语义决策先验,降低输入维度并注入任务相关知识,加速训练并提升策略可解释性。然而,如何将高层语义指导与连续低层控制对齐仍是难点。为此,引入一致性损失,促使VLM的语义计划与RL代理的控制输出保持一致,从而增强可解释性与训练稳定性。在CARLA模拟器上的实验表明,COVLM-RL在已训练环境中的成功率提升30%,在未见环境中的成功率提升50%,展现出强大的泛化能力。

原文摘要 · Abstract (English)

End-to-end autonomous driving frameworks face persistent challenges in generalization, training efficiency, and interpretability. While recent methods leverage Vision-Language Models (VLMs) through supervised learning on large-scale datasets to improve reasoning, they often lack robustness in novel scenarios. Conversely, reinforcement learning (RL)-based approaches enhance adaptability but remain data-inefficient and lack transparent decision-making. % contribution To address these limitations, we propose COVLM-RL, a novel end-to-end driving framework that integrates Critical Object-oriented (CO) reasoning with VLM-guided RL. Specifically, we design a Chain-of-Thought (CoT) prompting strategy that enables the VLM to reason over critical traffic elements and generate high-level semantic decisions, effectively transforming multi-view visual inputs into structured semantic decision priors. These priors reduce the input dimensionality and inject task-relevant knowledge into the RL loop, accelerating training and improving policy interpretability. However, bridging high-level semantic guidance with continuous low-level control remains non-trivial. To this end, we introduce a consistency loss that encourages alignment between the VLM's semantic plans and the RL agent's control outputs, enhancing interpretability and training stability. Experiments conducted in the CARLA simulator demonstrate that COVLM-RL significantly improves the success rate by 30\% in trained driving environments and by 50\% in previously unseen environments, highlighting its strong generalization capability.

自动驾驶视觉语言模型强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。