让视觉语言模型学会理解道路几何,提升复杂路况下的自动驾驶规划能力。
Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics

- 通过内化地图语义,增强视觉表征对道路结构的感知。
- 在单摄像头场景下达到92.1 PDMS,刷新单目VLA规划性能纪录。
- 无需高精地图,可直接部署于现有VLA系统,适配性强。
视觉-语言-动作(VLA)模型通过基础模型实现语义推理与长尾泛化,推动端到端自动驾驶发展。然而,在复杂驾驶环境中,仅依赖图像的表征难以捕捉与规划相关的道路几何与拓扑结构,限制了规划性能。本文提出Geo-VLA,一种即插即用框架,通过学习几何感知的视觉表征来增强VLA模型。训练阶段,Geo-VLA将地图语义内化以强化道路结构表征,推理时无需高精地图或额外车道信息。为此,我们构建了面向几何的问答数据集Geo-QA,通过对比学习与指令微调将道路几何知识注入视觉-语言表示。在NAVSIM v1上的实验表明,Geo-VLA在不同动作生成架构下均显著提升规划性能,达到92.1 PDMS,成为当前单摄像头VLA规划器的新基准。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments because image-only representations inadequately capture planning-relevant road geometry and topology. In this paper, we propose Geo-VLA, a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations. During training, Geo-VLA internalizes geometric map semantics to strengthen road-structure representations, while requiring no HD maps or additional lane information during inference. To support this approach, we introduce Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning. Experiments on NAVSIM v1 demonstrate that Geo-VLA consistently improves VLA planners with distinct action-generation architectures, achieving 92.1 PDMS and establishing a new state-of-the-art among single-camera VLA planners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。