构建首个面向行人骑行者事故的多模态推理评测基准
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
- 聚焦真实交通事故视频,设计多模态问答与密集描述任务
- 涵盖1000段视频、6000个问题、24000个选项,精细标注事故因果关系
- 揭示大模型在事故原因推理上仍存明显短板,适合安全驾驶研究者使用
保障行人和骑行者等弱势道路使用者(VRUs)的安全是自动驾驶系统的关键挑战,因其事故常导致严重或致命后果。尽管多模态大语言模型(MLLMs)在提升车辆场景理解与决策方面展现出潜力,但当前尚无标准化基准来量化评估其在涉及VRUs的高风险交通场景中的推理能力。为填补这一空白,我们提出VRU-Accident,一个大规模视觉-语言基准,用于评估MLLMs在高危交通场景中对VRU-车辆事故的理解能力。该基准包含1000段真实行车记录仪事故视频,标注了6000个多选问答对(覆盖6个安全关键类别,共24000个候选选项,3400个唯一答案),以及1000条密集场景描述。与以往工作不同,本基准明确聚焦于VRU-车辆事故,提供丰富的细粒度标注,涵盖事故的空间-时间动态与因果语义。为评估当前MLLMs水平,我们在多选问答与密集描述任务上对17个前沿模型进行了全面测试。结果表明,虽然这些模型在视觉基础属性识别上表现尚可,但在事故成因、类型及可预防性推理方面仍面临显著挑战。
原文摘要 · Abstract (English)
Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, is a critical challenge for autonomous driving systems, as crashes involving VRUs often result in severe or fatal consequences. While multimodal large language models (MLLMs) have shown promise in enhancing scene understanding and decision making in autonomous vehicles, there is currently no standardized benchmark to quantitatively evaluate their reasoning abilities in complex, safety-critical scenarios involving VRUs. To address this gap, we present VRU-Accident, a large-scale vision-language benchmark designed to evaluate MLLMs in high-risk traffic scenarios involving VRUs. VRU-Accident comprises 1K real-world dashcam accident videos, annotated with 6K multiple-choice question-answer pairs across six safety-critical categories (with 24K candidate options and 3.4K unique answer choices), as well as 1K dense scene descriptions. Unlike prior works, our benchmark focuses explicitly on VRU-vehicle accidents, providing rich, fine-grained annotations that capture both spatial-temporal dynamics and causal semantics of accidents. To assess the current landscape of MLLMs, we conduct a comprehensive evaluation of 17 state-of-the-art models on the multiple-choice VQA task and on the dense captioning task. Our findings reveal that while MLLMs perform reasonably well on visually grounded attributes, they face significant challenges in reasoning and describing accident causes, types, and preventability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。