让自动驾驶模型主动看图决策,提升可靠性与可解释性。
DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking
- 引入主动感知机制,遇不确定时调用视觉工具获取证据
- 3B参数下性能媲美GPT-5和人类驾驶,在长尾场景表现突出
- 融合文本与视觉双重推理,适合追求高可靠性的自动驾驶研发
视觉语言模型(VLM)推动了端到端自动驾驶的发展,展现出强大的高层行为规划推理能力。然而,现有方法多依赖被动感知与纯文本推理,难以在不确定性下主动获取关键视觉信息。为此,我们提出首个具备主动感知能力的自动驾驶代理DriveAgent-R1。在复杂场景中,它能主动调用工具进行视觉推理,使决策牢固基于视觉证据,从而提升可解释性与可靠性。同时,我们设计了一种受人类驾驶认知启发的混合思维框架,可根据场景复杂度自适应切换高效纯文本推理与稳健的工具增强视觉推理。该能力通过三阶段渐进式训练策略培养,核心为级联强化学习(Cascaded RL)。在包含长尾场景的Drive-Internal数据集及公开nuScenes数据集上的实验表明,仅用3B参数的DriveAgent-R1即达到与顶尖闭源模型(如GPT-5)相当的性能,并接近人类驾驶水平,同时保持部署友好性,为构建更智能的自动驾驶系统提供可行路径。
原文摘要 · Abstract (English)
The advent of Vision-Language Models (VLMs) has significantly advanced end-to-end autonomous driving, demonstrating powerful reasoning abilities for high-level behavior planning tasks. However, existing methods are often constrained by a passive perception paradigm, relying solely on text-based reasoning. This passivity restricts the model's capacity to actively seek crucial visual evidence when faced with uncertainty. To address this, we introduce DriveAgent-R1, the first autonomous driving agent capable of active perception for planning. In complex scenarios, DriveAgent-R1 proactively invokes tools to perform visual reasoning, firmly grounding its decisions in visual evidence, thereby enhancing both interpretability and reliability. Furthermore, we propose a hybrid thinking framework, inspired by human driver cognitive patterns, allowing the agent to adaptively switch between efficient text-only reasoning and robust tool-augmented visual reasoning based on scene complexity. This capability is cultivated through a three-stage progressive training strategy, featuring a core Cascaded Reinforcement Learning (Cascaded RL) phase. Extensive experiments on the Drive-Internal dataset, which is rich in long-tail scenarios, and the public nuScenes dataset show that, with only 3B parameters, DriveAgent-R1 achieves competitive performance comparable to top closed model systems such as GPT-5 and to human driving proficiency while remaining deployment-friendly, offering a proven path toward building more intelligent autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。