arXiv:2604.08457cs.CVcs.AI2026-04被引 4

构建道路事故视频基准,评估模型在真实场景中的推理能力。

CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning

论文配图:CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
图 1 · 摘自论文原文
  • 基于路侧摄像头数据构建250段事故视频的多选题评测集。
  • 8个顶尖视觉语言模型在因果与时间推理上表现不佳。
  • 适合研究自动驾驶协同感知与安全推理的学者使用。

协同自动驾驶需要从车辆和基础设施双视角理解交通场景。尽管视觉语言模型具备强大泛化推理能力,但现有基准多聚焦于车载视角,难以评估其在高危交通场景下的表现。为此,我们提出CrashSight,一个基于真实路侧摄像头数据的大规模视觉语言基准,涵盖250段事故视频,标注13,000个多项选择题,按两级分类体系组织:一级评估场景上下文与涉事方的视觉定位,二级考察碰撞机制、因果归因、时间演进及事故后结果等高级推理。我们对8个前沿视觉语言模型进行评测,发现尽管它们在场景描述上表现良好,但在安全关键场景中仍面临时间和因果推理瓶颈。本文详细分析失败案例,并探讨提升方向。该基准为协同自动驾驶中的基础设施辅助感知提供了标准化评估框架。完整数据集与代码已公开于https://mcgrche.github.io/crashsight。

原文摘要 · Abstract (English)

Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical traffic scenarios remains insufficiently evaluated due to the ego-vehicle focus of existing benchmarks. To bridge this gap, we present \textbf{CrashSight}, a large-scale vision-language benchmark for roadway crash understanding using real-world roadside camera data. The dataset comprises 250 crash videos, annotated with 13K multiple-choice question-answer pairs organized under a two-tier taxonomy. Tier 1 evaluates the visual grounding of scene context and involved parties, while Tier 2 probes higher-level reasoning, including crash mechanics, causal attribution, temporal progression, and post-crash outcomes. We benchmark 8 state-of-the-art VLMs and show that, despite strong scene description capabilities, current models struggle with temporal and causal reasoning in safety-critical scenarios. We provide a detailed analysis of failure scenarios and discuss directions for improving VLM crash understanding. The benchmark provides a standardized evaluation framework for infrastructure-assisted perception in cooperative autonomous driving. The CrashSight benchmark, including the full dataset and code, is accessible at https://mcgrche.github.io/crashsight.

事故理解视觉语言模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。