构建首个交通事故检测基准数据集,支持多种场景评估
ACCIDENT: A Benchmark Dataset for Vehicle Accident Detection from Traffic Surveillance Videos
- 构建2027段真实+2211段合成视频,标注时间、位置与碰撞类型
- 在零样本和跨域设置下,现有方法性能普遍低于人类水平
- 适合做交通事故检测、视频理解与少样本学习的研究者
我们提出ACCIDENT,一个用于监控摄像头视频中交通事故检测的基准数据集,旨在评估模型在有监督(同分布与异分布)及零样本设置下的表现,涵盖数据丰富与数据稀缺两种场景。该数据集包含2,027段真实视频与2,211段合成视频,均标注了事故发生时间、空间位置及高层次碰撞类型。定义三项核心任务:(i)事故时间定位,(ii)空间定位,(iii)碰撞类型分类。每项任务采用定制化评估指标,以应对监控视频中固有的不确定性和模糊性。除数据集外,还提供多样化基线模型,包括启发式、运动感知及视觉-语言方法,并验证了该基准的挑战性。数据集可访问:https://accidentbench.github.io
原文摘要 · Abstract (English)
We introduce ACCIDENT, a benchmark dataset for traffic accident detection in CCTV footage, designed to evaluate models in supervised (IID and OOD) and zero-shot settings, reflecting both data-rich and data-scarce scenarios. The benchmark consists of a curated set of 2,027 real and 2,211 synthetic clips annotated with the accident time, spatial location, and high-level collision type. We define three core tasks: (i) temporal localization of the accident, (ii) its spatial localization, and (iii) collision type classification. Each task is evaluated using custom metrics that account for the uncertainty and ambiguity inherent in CCTV footage. In addition to the benchmark, we provide a diverse set of baselines, including heuristic, motion-aware, and vision-language approaches, and show that ACCIDENT is challenging. You can access the ACCIDENT at: https://accidentbench.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。