arXiv:2601.04714cs.AI2026-01被引 2

用思维链引导强化学习,让自动驾驶决策更像人。

ThinkDrive: Chain-of-Thought Guided Progressive Reinforcement Learning Fine-Tuning for Autonomous Driving

  • 用思维链逐步增强驾驶决策的透明性与逻辑性
  • 在公开数据集上超越强基线1.45%~3.28%
  • 适合关注可解释自动驾驶的开发者与研究者

随着大语言模型技术的快速发展,其在自动驾驶领域的应用日益广泛。然而,现有方法存在推理不结构化、泛化能力差以及与人类驾驶意图不一致等问题。尽管思维链(CoT)推理提升了决策透明度,但传统监督微调(SFT)未能充分发挥其潜力,而强化学习(RL)方法则面临训练不稳定和推理深度不足的挑战。本文提出ThinkDrive,一种基于思维链引导的渐进式强化学习微调框架,通过显式推理与难度感知的自适应策略优化相结合,实现更可靠的驾驶决策。该方法采用两阶段训练:首先使用含思维链解释的SFT进行预训练;随后引入难度感知的自适应策略优化器,根据样本复杂度动态调整学习强度。我们在公开数据集上评估了该方法,结果表明,ThinkDrive在exam、easy-exam、accuracy三项指标上分别优于强基线1.45%、1.95%和1.01%。此外,一个20亿参数的模型经本方法训练后,在exam指标上超越了更大的GPT-4o模型3.28%。

原文摘要 · Abstract (English)

With the rapid advancement of large language models (LLMs) technologies, their application in the domain of autonomous driving has become increasingly widespread. However, existing methods suffer from unstructured reasoning, poor generalization, and misalignment with human driving intent. While Chain-of-Thought (CoT) reasoning enhances decision transparency, conventional supervised fine-tuning (SFT) fails to fully exploit its potential, and reinforcement learning (RL) approaches face instability and suboptimal reasoning depth. We propose ThinkDrive, a CoT guided progressive RL fine-tuning framework for autonomous driving that synergizes explicit reasoning with difficulty-aware adaptive policy optimization. Our method employs a two-stage training strategy. First, we perform SFT using CoT explanations. Then, we apply progressive RL with a difficulty-aware adaptive policy optimizer that dynamically adjusts learning intensity based on sample complexity. We evaluate our approach on a public dataset. The results show that ThinkDrive outperforms strong RL baselines by 1.45%, 1.95%, and 1.01% on exam, easy-exam, and accuracy, respectively. Moreover, a 2B-parameter model trained with our method surpasses the much larger GPT-4o by 3.28% on the exam metric.

自动驾驶思维链强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。