arXiv:2412.20042cs.CV2024-12

DAVE数据集聚焦复杂交通中弱势道路使用者,提升感知算法真实场景适应性。

DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments

  • 构建包含16类对象与16种行为的高密度标注视频数据集
  • 弱势道路使用者占比41.13%,显著高于Waymo的23.71%
  • 适用于追踪、动作定位等任务,适合自动驾驶感知研究

现有交通视频数据集(如Waymo)多聚焦西方交通场景,难以覆盖全球复杂环境。针对亚洲地区更复杂的交通状况,本文提出DAVE数据集,专门用于评估感知方法在复杂不可预测环境中的表现。该数据集手动标注了16类不同主体(含动物、人类、车辆等)和16种复杂罕见行为(如插队、蛇形行驶、掉头等),要求高推理能力。整体标注超过1300万帧边界框,其中160多万个同时标注了主体身份与行为信息。视频采集涵盖多种天气、时段、道路类型及交通密度。DAVE可用于追踪、检测、时空动作定位、图文瞬间检索和多标签视频动作识别等任务。由于弱势道路使用者(VRUs)占总实例的41.13%(远高于Waymo的23.71%),对事故预防至关重要。实验表明,现有方法在该数据集上性能显著下降,凸显其对推动未来视频识别研究的价值。

原文摘要 · Abstract (English)

Most existing traffic video datasets including Waymo are structured, focusing predominantly on Western traffic, which hinders global applicability. Specifically, most Asian scenarios are far more complex, involving numerous objects with distinct motions and behaviors. Addressing this gap, we present a new dataset, DAVE, designed for evaluating perception methods with high representation of Vulnerable Road Users (VRUs: e.g. pedestrians, animals, motorbikes, and bicycles) in complex and unpredictable environments. DAVE is a manually annotated dataset encompassing 16 diverse actor categories (spanning animals, humans, vehicles, etc.) and 16 action types (complex and rare cases like cut-ins, zigzag movement, U-turn, etc.), which require high reasoning ability. DAVE densely annotates over 13 million bounding boxes (bboxes) actors with identification, and more than 1.6 million boxes are annotated with both actor identification and action/behavior details. The videos within DAVE are collected based on a broad spectrum of factors, such as weather conditions, the time of day, road scenarios, and traffic density. DAVE can benchmark video tasks like Tracking, Detection, Spatiotemporal Action Localization, Language-Visual Moment retrieval, and Multi-label Video Action Recognition. Given the critical importance of accurately identifying VRUs to prevent accidents and ensure road safety, in DAVE, vulnerable road users constitute 41.13% of instances, compared to 23.71% in Waymo. DAVE provides an invaluable resource for the development of more sensitive and accurate visual perception algorithms in the complex real world. Our experiments show that existing methods suffer degradation in performance when evaluated on DAVE, highlighting its benefit for future video recognition research.

自动驾驶视觉感知数据集弱势道路使用者

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。