arXiv:2601.05640cs.CV2026-01被引 28

用分层认知结构让通用视觉模型更懂开车。

SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving

  • 将驾驶理解分解为场景-车辆-目标三层结构,模仿人类驾驶思维。
  • 在NAVSIM基准上,相机仅输入下达到最佳表现,PDMS和EPDMS双指标领先。
  • 适合研究端到端自动驾驶与多模态模型融合的开发者参考。

近期端到端自动驾驶方法利用视觉语言模型(VLM)提升复杂场景下的规划能力,但这些通用模型缺乏对三维时空中驾驶特有推理的理解。在应用于自动驾驶时,它们难以构建结构化的时空表征,无法有效捕捉几何关系、场景上下文和运动模式,影响安全轨迹规划。为此,我们提出SGDrive框架,基于预训练的VLM主干,显式构建以驾驶知识层级为核心的表征学习体系。该框架将驾驶理解划分为场景-代理-目标的分层结构,模拟人类驾驶认知:先感知整体环境(场景上下文),再关注关键交通参与者及其行为,最后制定短期目标并执行动作。这一分层设计提供了通用VLM所欠缺的结构化时空表示,将多层级信息融合为紧凑而全面的格式用于轨迹规划。在NAVSIM基准上的大量实验表明,SGDrive在仅使用摄像头输入的方法中,在PDMS和EPDMS两个指标上均达到当前最优性能,验证了分层知识结构在适配通用VLM至自动驾驶任务中的有效性。

原文摘要 · Abstract (English)

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized understanding of driving-specific reasoning in 3D space and time. When applied to autonomous driving, these models struggle to establish structured spatial-temporal representations that capture geometric relationships, scene context, and motion patterns critical for safe trajectory planning. To address these limitations, we propose SGDrive, a novel framework that explicitly structures the VLM's representation learning around driving-specific knowledge hierarchies. Built upon a pre-trained VLM backbone, SGDrive decomposes driving understanding into a scene-agent-goal hierarchy that mirrors human driving cognition: drivers first perceive the overall environment (scene context), then attend to safety-critical agents and their behaviors, and finally formulate short-term goals before executing actions. This hierarchical decomposition provides the structured spatial-temporal representation that generalist VLMs lack, integrating multi-level information into a compact yet comprehensive format for trajectory planning. Extensive experiments on the NAVSIM benchmark demonstrate that SGDrive achieves state-of-the-art performance among camera-only methods on both PDMS and EPDMS, validating the effectiveness of hierarchical knowledge structuring for adapting generalist VLMs to autonomous driving.

自动驾驶视觉语言模型分层认知端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。