首个面向自动驾驶的视觉语言模型空间理解评测基准
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
- 基于NuScenes数据集构建3D场景图与问答对自动生成流水线
- 发现增强空间能力的模型在定性问答中表现更好但定量表现不佳
- 适合评估视觉语言模型在自动驾驶中的空间推理能力
视觉语言模型(VLMs)在自动驾驶任务中展现出巨大潜力,但其空间理解与推理——自动驾驶的关键能力——仍存在明显局限。现有基准未系统评估VLM在驾驶场景下的空间推理能力。为此,我们提出NuScenes-SpatialQA,首个基于真实标注的大规模问答基准,专门用于评估VLM在自动驾驶中的空间理解与推理能力。该基准基于NuScenes数据集,通过自动化3D场景图生成与问答生成流程构建。它从多个维度系统评估VLM在空间理解与推理方面的能力。我们在包括通用与空间增强型在内的多种VLM上进行了广泛实验,首次全面评估其在自动驾驶中的空间能力。令人意外的是,空间增强型VLM在定性问答中表现更优,但在定量问答中并无竞争力。总体而言,当前VLM在空间理解与推理方面仍面临显著挑战。
原文摘要 · Abstract (English)
Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhibit significant limitations. Notably, none of the existing benchmarks systematically evaluate VLMs' spatial reasoning capabilities in driving scenarios. To fill this gap, we propose NuScenes-SpatialQA, the first large-scale ground-truth-based Question-Answer (QA) benchmark specifically designed to evaluate the spatial understanding and reasoning capabilities of VLMs in autonomous driving. Built upon the NuScenes dataset, the benchmark is constructed through an automated 3D scene graph generation pipeline and a QA generation pipeline. The benchmark systematically evaluates VLMs' performance in both spatial understanding and reasoning across multiple dimensions. Using this benchmark, we conduct extensive experiments on diverse VLMs, including both general and spatial-enhanced models, providing the first comprehensive evaluation of their spatial capabilities in autonomous driving. Surprisingly, the experimental results show that the spatial-enhanced VLM outperforms in qualitative QA but does not demonstrate competitiveness in quantitative QA. In general, VLMs still face considerable challenges in spatial understanding and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。