让自动驾驶模型会思考、能自省,提升决策可靠性。
AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
- 用思维链+自省机制构建决策逻辑,提升可解释性。
- 在nuScenes和Waymo数据集上达到当前最佳性能。
- 适合关注自动驾驶决策可信度的研究者与工程师。
自动驾驶中的视觉-语言-动作(VLA)模型虽已展现巨大潜力,但其决策过程的可解释性与动作序列的合理性仍待深入探索。为此,我们提出AutoDrive-R²框架,通过思维链(CoT)处理与强化学习(RL)增强模型的推理与自省能力。首先构建了包含6000条样本的nuScenesR²-6K CoT数据集,采用四步逻辑链并引入自省验证,建立输入信息与输出轨迹间的认知桥梁。其次,在强化学习阶段,采用组相对策略优化(GRPO)算法,在融合空间对齐、车辆动力学与时间平滑性的物理约束奖励框架下,最大化推理与自省效果,确保轨迹规划的可靠与真实。在nuScenes与Waymo数据集上的广泛评估表明,该方法实现当前最优性能并具备强泛化能力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models in autonomous driving systems have recently demonstrated transformative potential by integrating multimodal perception with decision-making capabilities. However, the interpretability and coherence of the decision process and the plausibility of action sequences remain largely underexplored. To address these issues, we propose AutoDrive-R$^2$, a novel VLA framework that enhances both reasoning and self-reflection capabilities of autonomous driving systems through chain-of-thought (CoT) processing and reinforcement learning (RL). Specifically, we first propose an innovative CoT dataset named nuScenesR$^2$-6K for supervised fine-tuning, which effectively builds cognitive bridges between input information and output trajectories through a four-step logical chain with self-reflection for validation. Moreover, to maximize both reasoning and self-reflection during the RL stage, we further employ the Group Relative Policy Optimization (GRPO) algorithm within a physics-grounded reward framework that incorporates spatial alignment, vehicle dynamic, and temporal smoothness criteria to ensure reliable and realistic trajectory planning. Extensive evaluation results across both nuScenes and Waymo datasets demonstrates the state-of-the-art performance and robust generalization capacity of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。