arXiv:2608.09591cs.ROcs.CV2026-08

让自动驾驶模型根据关键因素自适应推理,提升规划质量。

FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving

论文配图:FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving
图 1 · 摘自论文原文
  • 基于关键因素构建可动态调整的推理链,融合空间物理证据
  • 通过奖励引导搜索发现更优规划路径,显著提升轨迹生成质量
  • 适合追求高精度端到端自动驾驶系统的研究人员与开发者

视觉语言模型(VLMs)已推动场景理解并实现端到端自动驾驶中的显式推理。然而,现有方法在规划推理中对空间-物理证据的整合不足,且推理适应性粗糙,难以满足场景特定需求。此外,后训练阶段的推理路径优化仍缺乏探索。为此,我们提出FactorDrive,一种由规划关键因素(PCFs)驱动的自适应多步推理端到端自动驾驶框架。首先进行大规模驾驶领域指令微调以建立基础驾驶知识。在此基础上,构建PCF-CoT——一个将规划推理锚定于轨迹相关空间-物理证据、围绕场景特异性PCFs组织的思维链数据集,使推理路径的组合与深度能随不同规划需求自适应调整。进一步提出质量搜索引导的组相对策略优化(QS-GRPO),利用轨迹级规划奖励指导蒙特卡洛树搜索(MCTS)以发现更高规划质量的推理路径,并用所得结果通过组相对策略优化(GRPO)更新策略,从而提升轨迹规划性能。在开环(nuScenes)与闭环导向(NAVSIM)基准上的大量实验表明,FactorDrive实现了顶尖的规划表现。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of scene-specific planning demands. Furthermore, reasoning-path optimization for higher planning quality remains largely unexplored in autonomous-driving post-training. To address these limitations, we propose FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs). We first perform large-scale driving-domain instruction tuning to establish foundational driving knowledge. Building on this foundation, we construct PCF-CoT, a chain-of-thought (CoT) dataset that grounds planning reasoning in trajectory-relevant spatial-physical evidence and organizes reasoning around scene-specific PCFs, enabling the composition and depth of reasoning paths to adapt to different planning demands. We further introduce Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which guides Monte Carlo Tree Search (MCTS) with trajectory-level planning rewards to discover reasoning paths with higher planning quality and uses the resulting responses to optimize the policy through GRPO, thereby improving trajectory planning performance. Extensive experiments on both open-loop (nuScenes) and closed-loop-oriented (NAVSIM) benchmarks demonstrate that FactorDrive achieves state-of-the-art planning performance.

自动驾驶推理链规划优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。