评测视觉语言模型在连续驾驶场景中的表现,发现其准确率仅57%。
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study

- 构建VENUSS框架,系统分析输入配置对模型性能的影响
- 25+模型测试中最高准确率57%,低于人类的65%
- 揭示模型在车辆动态与时间关系理解上的明显短板
视觉语言模型(VLMs)被越来越多用于自动驾驶任务,但其在连续驾驶场景中的表现仍缺乏系统评估,尤其不清楚输入配置如何影响其能力。我们提出VENUSS(VLM Evaluation oN Understanding Sequential Scenes),一个针对连续驾驶场景的系统性敏感性分析框架,为未来研究建立基准。基于现有数据集,VENUSS从驾驶视频中提取时间序列,并生成结构化评估,覆盖自定义类别。通过对比25+现有VLMs在2,600+场景中的表现,我们发现即使顶级模型也仅达57%准确率,未达到人类在相似约束下的65%表现,暴露出显著的能力差距。分析显示,VLMs在静态物体检测上表现良好,但在车辆动态和时间关系理解上存在困难。VENUSS首次系统分析了输入图像配置(分辨率、帧数、时间间隔、空间布局、呈现方式)对连续驾驶场景理解性能的影响。补充材料见https://TUM-AVS.github.io/VENUSS/。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poorly characterized, particularly regarding how input configurations affect their capabilities. We introduce VENUSS (VLM Evaluation oN Understanding Sequential Scenes), a framework for systematic sensitivity analysis of VLM performance on sequential driving scenes, establishing baselines for future research. Building upon existing datasets, VENUSS extracts temporal sequences from driving videos, and generates structured evaluations across custom categories. By comparing 25+ existing VLMs across 2,600+ scenarios, we reveal how even top models achieve only 57% accuracy, not matching human performance under similar constraints (65%) and exposing significant capability gaps. Our analysis shows that VLMs excel with static object detection but struggle with understanding vehicle dynamics and temporal relations. VENUSS offers the first systematic sensitivity analysis of VLMs focused on how input image configurations - resolution, frame count, temporal intervals, spatial layouts, and presentation modes - affect performance on sequential driving scenes. Supplementary material available at https://TUM-AVS.github.io/VENUSS/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。