arXiv:2506.06218cs.CV2025-06NeurIPS被引 20

构建多视角交通场景基准,评估自动驾驶大模型的时空推理能力

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

  • 从数据中自动提取交通场景,人工验证后生成选择题评测模型
  • 覆盖43种场景、971道题目,聚焦车辆行为与交互的时空理解
  • 适合研究自动驾驶视觉语言模型的科研人员和工程师

我们提出STSBench,一个基于场景的框架,用于评估自动驾驶视觉语言模型(VLMs)的综合理解能力。该框架通过真实标注自动挖掘任意数据集中的预定义交通场景,提供直观的人工验证界面,并生成多选题用于模型评估。应用于NuScenes数据集,我们构建了首个基于3D感知的时空推理基准STSnu。现有基准多针对单视角图像或视频的现成/微调模型,侧重物体识别、密集描述等语义任务。而STSnu评估端到端驾驶场景下的驾驶专家型VLMs,处理多视角摄像头或激光雷达视频,重点考察其对自身车辆动作及交通参与者复杂交互的推理能力,这是自动驾驶的关键需求。基准包含43种多样化场景,覆盖多视角与多帧,生成971道经人工验证的多选题。全面评估揭示现有模型在复杂环境中理解基础交通动态方面存在显著缺陷,凸显亟需显式建模时空推理的架构改进。通过填补时空评估的核心空白,STSBench推动更鲁棒、可解释的自动驾驶VLMs发展。

原文摘要 · Abstract (English)

We introduce STSBench, a scenario-based framework to benchmark the holistic understanding of vision-language models (VLMs) for autonomous driving. The framework automatically mines pre-defined traffic scenarios from any dataset using ground-truth annotations, provides an intuitive user interface for efficient human verification, and generates multiple-choice questions for model evaluation. Applied to the NuScenes dataset, we present STSnu, the first benchmark that evaluates the spatio-temporal reasoning capabilities of VLMs based on comprehensive 3D perception. Existing benchmarks typically target off-the-shelf or fine-tuned VLMs for images or videos from a single viewpoint and focus on semantic tasks such as object recognition, dense captioning, risk assessment, or scene understanding. In contrast, STSnu evaluates driving expert VLMs for end-to-end driving, operating on videos from multi-view cameras or LiDAR. It specifically assesses their ability to reason about both ego-vehicle actions and complex interactions among traffic participants, a crucial capability for autonomous vehicles. The benchmark features 43 diverse scenarios spanning multiple views and frames, resulting in 971 human-verified multiple-choice questions. A thorough evaluation uncovers critical shortcomings in existing models' ability to reason about fundamental traffic dynamics in complex environments. These findings highlight the urgent need for architectural advances that explicitly model spatio-temporal reasoning. By addressing a core gap in spatio-temporal evaluation, STSBench enables the development of more robust and explainable VLMs for autonomous driving.

自动驾驶多模态时空推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。