arXiv:2509.16654cs.CV2025-09

评测视觉语言模型在车道拓扑理解上的能力,发现其空间推理仍不成熟。

Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?

  • 将多视角图像融合为鸟瞰图车道,构建四类拓扑诊断问答任务。
  • 前沿闭源模型如GPT-4o在部分任务中准确率仅67.8%,人类可轻松解答。
  • 开源模型(30B规模)表现更差,模型大小与推理长度影响其性能。

视觉语言模型(VLMs)在多模态推理方面取得显著进展,但在自动驾驶中的应用仍有限,尤其对道路拓扑的理解关注不足。本文系统评估了VLM在道路拓扑理解方面的能力:将多视角图像投影并融合为统一地平面坐标系下的鸟瞰图车道,基于此构建四类拓扑相关的诊断型VQA任务,涵盖空间拓扑推理的关键要素。大量实验表明,尽管前沿闭源模型(如GPT-4o)在某些任务中表现较好,但在人类可轻松回答的空间问题上仍存在明显短板(如向量分类任务准确率仅67.8%)。此外,开源模型(即使达到30B规模)表现显著不足。结果表明,空间推理仍是当前VLM的核心瓶颈。我们还发现模型能力与模型规模、推理令牌长度及示例数量正相关,为未来研究指明方向。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have recently shown remarkable progress in multimodal reasoning, yet their applications in autonomous driving remain limited. In particular, the ability to understand road topology, a key requirement for safe navigation, has received relatively little attention. While some recent works have begun to explore VLMs in driving contexts, their performance on topology reasoning is far from satisfactory. In this work, we systematically evaluate VLMs' capabilities in road topology understanding. Specifically, multi-view images are projected into unified ground-plane coordinate system and fused into bird's-eye-view (BEV) lanes. Based on these BEV lanes, we formulate four topology-related diagnostic VQA tasks, which together capture essential components of spatial topology reasoning. Through extensive evaluation, we find that while frontier closed-source models (e.g., GPT-4o) achieve relatively high accuracy in some tasks, they still fail in some spatial questions that humans can answer (e.g., GPT-4o achieve only 67.8% in vector, a two-class classification problem). Furthermore, we find open-source VLMs, even at 30B scale, struggle significantly. These results indicate that spatial reasoning remains a fundamental bottleneck for current VLMs. We also find that the model's capability is positively correlated with model size, length of reasoning tokens and shots provided as examples, showing direction for future research.

视觉语言模型自动驾驶空间推理拓扑理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。