arXiv:2512.01830cs.CV2025-12被引 10

用大模型当裁判,让自动驾驶全程自我优化推理与规划。

OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic

  • 以大模型为裁判,对开放性驾驶推理进行奖励建模。
  • 端到端强化微调使推理与路径规划性能全面领先。
  • 适合研究智能驾驶认知推理与强化学习融合的学者。

近期两阶段微调策略(如通过监督微调获取驾驶知识,再通过强化微调提升决策与规划能力)在知识驱动型自动驾驶中展现出巨大潜力。然而,监督微调的学习特性限制了推理泛化能力,制约了驾驶性能的充分发挥。同时,现有强化微调方法多局限于下游任务,因场景理解是开放性问题,难以量化奖励。为此,我们提出OpenREAD,一种基于视觉语言模型的端到端自动驾驶框架,支持从高层推理到低层轨迹规划的全流程强化微调。具体而言,我们在开源驾驶相关知识数据集上构建大规模思维链(CoT)标注,并利用强大的Qwen3大语言模型作为强化微调中的裁判,量化开放性问题的推理质量。大量实验表明,联合端到端强化微调显著提升上下游任务表现,使OpenREAD在推理与规划基准上达到当前最优水平。

原文摘要 · Abstract (English)

Recently, two-stage fine-tuning strategies, e.g., acquiring essential driving knowledge through supervised fine-tuning (SFT) and further enhancing decision-making and planning via reinforcement fine-tuning (RFT), have shown strong potential in advancing the knowledge-driven autonomous driving (AD) paradigm. However, the learning nature of SFT still limits the generalization of reasoning, thereby constraining the full potential of driving performance. Meanwhile, current RFT approaches are primarily applied to downstream tasks, since scene understanding is an open-ended problem where corresponding rewards are difficult to quantify. To address these limitations, we propose OpenREAD, an OPEN-ended REasoning reinforced vision-language model (VLM)-based autonomous driving (AD) framework that enables end-to-end RFT across the full spectrum from high-level reasoning to low-level trajectory planning. Specifically, we begin by constructing large-scale Chain-of-Thought (CoT) annotations on open-source driving-related knowledge datasets, and employ the powerful Qwen3 large language model (LLM) as the critic in RFT to quantify reasoning quality for open-ended questions during reward modeling. Extensive experiments confirm that joint end-to-end RFT yields substantial improvements in both upstream and downstream tasks, enabling OpenREAD to achieve state-of-the-art performance on reasoning and planning benchmarks.

自动驾驶大模型强化学习推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。