arXiv:2501.08168cs.AI2025-01被引 20

LeapVAD用双思维模型让自动驾驶更像人,能从经验中持续学习。

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking

  • 模仿人类注意力与双系统思维,聚焦关键交通元素。
  • 在CARLA和DriveArena上表现优于纯摄像头方案,少数据也有效。
  • 支持反思纠错与持续学习,适合需要长期进化的自动驾驶系统。

尽管自动驾驶技术取得显著进展,数据驱动方法在复杂场景中仍因推理能力有限而受限。知识驱动系统随视觉语言模型普及得到大幅提升。本文提出LeapVAD,基于认知感知与双过程思维的新方法。通过模拟人类注意力机制,识别并聚焦影响决策的关键交通元素,结合外观、运动模式与风险属性进行综合表征,实现更有效的环境表达并简化决策流程。引入创新的双过程决策模块,模拟人类驾驶学习过程:分析型过程(系统-II)通过逻辑推理积累经验,启发式过程(系统-I)则通过微调与少量样本学习优化知识。系统还具备反思机制与动态记忆库,在闭环环境中可从过往错误中学习,持续提升性能。为提高效率,设计场景编码网络,生成紧凑场景表示以快速检索相关驾驶经验。在CARLA与DriveArena两个主流仿真平台上的广泛评估表明,即使训练数据有限,LeapVAD仍显著优于纯摄像头方法。全面消融实验进一步验证其在持续学习与领域自适应中的有效性。

原文摘要 · Abstract (English)

While autonomous driving technology has made remarkable strides, data-driven approaches still struggle with complex scenarios due to their limited reasoning capabilities. Meanwhile, knowledge-driven autonomous driving systems have evolved considerably with the popularization of visual language models. In this paper, we propose LeapVAD, a novel method based on cognitive perception and dual-process thinking. Our approach implements a human-attentional mechanism to identify and focus on critical traffic elements that influence driving decisions. By characterizing these objects through comprehensive attributes - including appearance, motion patterns, and associated risks - LeapVAD achieves more effective environmental representation and streamlines the decision-making process. Furthermore, LeapVAD incorporates an innovative dual-process decision-making module miming the human-driving learning process. The system consists of an Analytic Process (System-II) that accumulates driving experience through logical reasoning and a Heuristic Process (System-I) that refines this knowledge via fine-tuning and few-shot learning. LeapVAD also includes reflective mechanisms and a growing memory bank, enabling it to learn from past mistakes and continuously improve its performance in a closed-loop environment. To enhance efficiency, we develop a scene encoder network that generates compact scene representations for rapid retrieval of relevant driving experiences. Extensive evaluations conducted on two leading autonomous driving simulators, CARLA and DriveArena, demonstrate that LeapVAD achieves superior performance compared to camera-only approaches despite limited training data. Comprehensive ablation studies further emphasize its effectiveness in continuous learning and domain adaptation. Project page: https://pjlab-adg.github.io/LeapVAD/.

自动驾驶双过程思维持续学习认知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。