arXiv:2605.23176cs.CV2026-05

构建动态驾驶场景多视角推理基准,揭示当前模型在时空理解上的巨大短板。

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

论文配图:DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
图 1 · 摘自论文原文
  • 基于动态多关系场景图生成跨视角问答对,强制模型进行真实时空推理。
  • 15个主流视觉语言模型平均落后人类28.4分,场景构建能力是主要瓶颈。
  • 显式鸟瞰图定位显著提升性能,语言提示 alone 不足应对复杂场景。

自动驾驶中的时空智能要求智能体将多视角观测整合为连贯的场景表征,维持物体在视角和时间上的连续性,并推理空间关系、交互与未来动态。然而现有自动驾驶视觉语言基准大多聚焦单视角、静态、自中心或单一来源的问答,难以评估模型是否真正具备动态场景构建与推理能力。我们提出DriveSpatial,涵盖20个大规模数据集的15.6万个经人工验证的问答对,覆盖五大类任务。该基准评估四大能力:认知场景构建、多视角关系理解、时间推理与泛化能力。不同于以往基准,DriveSpatial基于动态多关系场景图生成,编码物体状态、空间关系、交互、相机可见性及时间对应,使问答对强制实现真正的跨视角与时空推理。评估15个代表性视觉语言模型显示显著的人机差距:最强模型落后人类28.4分,其中认知场景构建成为关键瓶颈。进一步诊断表明,仅靠语言提示不足,而显式鸟瞰图(BEV)定位能持续提升性能。结果表明,当前视觉语言模型缺乏可靠时空驾驶智能所需的场景构建能力。DriveSpatial及其构建流程将公开发布,以支持后续研究。

原文摘要 · Abstract (English)

Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations, interactions, and future dynamics. However, existing AD vision-language benchmarks largely focus on single-view, static, ego-centric, or single-source question answering, leaving it unclear whether current Vision-Language Models (VLMs) can truly construct and reason over dynamic driving scenes. We introduce DriveSpatial, a benchmark of 15.6K human-verified QA pairs across 20 tasks from five large-scale AD datasets. DriveSpatial evaluates four abilities: Cognitive Scene Construction, Multi-view Relational Understanding, Temporal Reasoning, and Generalization. Unlike prior benchmarks, DriveSpatial is generated from a dynamic multi-relational scene graph that encodes object states, spatial relations, interactions, camera visibility, and temporal correspondences, enabling QA pairs that enforce genuine cross-view and spatiotemporal reasoning. Evaluating 15 representative VLMs reveals a substantial human-model gap: the strongest model trails humans by 28.4 points, with Cognitive Scene Construction emerging as the key bottleneck. Further diagnostics show that language-only prompting is insufficient, while explicit BEV grounding consistently improves performance. These results suggest that current VLMs lack the scene-construction ability needed for reliable spatiotemporal driving intelligence. DriveSpatial and its construction pipeline will be released to support future research.

自动驾驶视觉语言模型时空推理多视角理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。