构建事故与多场景安全任务基准,评估模型真实世界推理能力
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- 融合交通事故与空海安全场景,覆盖时空与意图推理任务
- 2000段视频+1.9万问答对,最难点任务顶尖模型仅18%准确率
- 适合研究多模态模型在真实安全场景下的可靠性与鲁棒性
多模态模型的快速发展亟需在安全关键、动态真实场景中严格评估其理解与推理能力。我们提出 AccidentBench,一个大规模基准,结合车辆事故场景与空海等安全关键领域,强调空间与时间推理(如导航、方位判断、多车运动)。该基准包含约2000段视频和超过19000个由人工标注的问题-答案对,涵盖短/中/长视频及易/中/难不同难度等级。任务系统性地检验时间、空间与意图理解与推理能力。通过统一事故导向交通场景与更广泛的空海安全场景,AccidentBench 提供了一个全面、物理基础坚实的测试平台,用于评估模型在真实世界变异性下的表现。对先进模型(如 Gemini-2.5 Pro 与 GPT-5)的评估显示,即使最强模型在最困难任务和最长视频上也仅达到约18%准确率,揭示了当前模型在真实世界时空与意图推理方面的显著差距。AccidentBench 旨在暴露这些关键缺陷,并推动更安全、更鲁棒、更契合现实安全挑战的多模态模型发展。代码与数据集已开源:https://github.com/SafeRL-Lab/AccidentBench
原文摘要 · Abstract (English)
Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench, a large-scale benchmark that combines vehicle accident scenarios with Beyond domains, safety-critical settings in air and water that emphasize spatial and temporal reasoning (e.g., navigation, orientation, multi-vehicle motion). The benchmark contains approximately 2000 videos and over 19000 human-annotated question--answer pairs spanning multiple video lengths (short/medium/long) and difficulty levels (easy/medium/hard). Tasks systematically probe core capabilities: temporal, spatial, and intent understanding and reasoning. By unifying accident-centric traffic scenes with broader safety-critical scenarios in air and water, AccidentBench offers a comprehensive, physically grounded testbed for evaluating models under real-world variability. Evaluations of state-of-the-art models (e.g., Gemini-2.5 Pro and GPT-5) show that even the strongest models achieve only about 18% accuracy on the hardest tasks and longest videos, revealing substantial gaps in real-world temporal, spatial, and intent reasoning. AccidentBench is designed to expose these critical gaps and drive the development of multimodal models that are safer, more robust, and better aligned with real-world safety-critical challenges. The code and dataset are available at: https://github.com/SafeRL-Lab/AccidentBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。