arXiv:2503.13047cs.CV2025-03被引 14

让自动驾驶更像人:用双重视角理解路况,提升决策能力

InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving

  • 构建显性与隐性双重场景表征,模仿人类注意力与推理方式
  • 引入思维链指令与轻量专家模块,在不增加参数的情况下注入驾驶知识
  • 基于扩散模型生成轨迹,适合追求高鲁棒性的自动驾驶研发者

传统端到端自动驾驶依赖显式的全局场景表征,包括3D目标检测、在线建图和运动预测。相比之下,人类驾驶员会聚焦任务相关区域,并隐式理解整体交通上下文。受此启发,我们提出轻量级端到端自动驾驶框架InsightDrive。该框架引入一种洞察式场景表征,联合建模以注意力为中心的显式表征与以推理为中心的隐式表征,使场景理解更贴近人类认知模式。为此,我们采用思维链(CoT)指令模拟驾驶认知,并设计任务级混合专家(MoE)适配器,在几乎不增加参数的前提下将知识注入模型。进一步地,规划器同时基于显性和隐性表征进行条件判断,并采用基于扩散的生成策略,实现稳健的轨迹预测与决策。整个框架建立了一个知识蒸馏流程,将人类驾驶知识传递至大语言模型,再迁移至车载模型。在nuScenes和Navsim基准上的大量实验表明,InsightDrive显著优于传统场景表征方法。

原文摘要 · Abstract (English)

Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion prediction. In contrast, human drivers selectively attend to task-relevant regions and implicitly reason over the broader traffic context. Motivated by this observation, we introduce a lightweight end-to-end autonomous driving framework, InsightDrive. Unlike approaches that directly embed large language models (LLMs), InsightDrive introduces an Insight scene representation that jointly models attention-centric explicit scene representation and reasoning-centric implicit scene representation, so that scene understanding aligns more closely with human cognitive patterns for trajectory planning. To this end, we employ Chain-of-Thought (CoT) instructions to model human driving cognition and design a task-level Mixture-of-Experts (MoE) adapter that injects this knowledge into the autonomous driving model at negligible parameter cost. We further condition the planner on both explicit and implicit scene representations and employ a diffusion-based generative policy, which produces robust trajectory predictions and decisions. The overall framework establishes a knowledge distillation pipeline that transfers human driving knowledge to LLMs and subsequently to onboard models. Extensive experiments on the nuScenes and Navsim benchmarks demonstrate that InsightDrive achieves significant improvements over conventional scene representation approaches.

自动驾驶场景表征扩散模型人类认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。