构建全球连续城市行车记录视频数据集,支持真实驾驶场景分析。
A global dataset of continuous urban dashcam driving
- 从公开YouTube视频中筛选并标注51,753段连续城市行车片段
- 覆盖238个国家地区、超2万小时视频,含昼夜与车辆类型标签
- 提供自动检测结果与多目标跟踪,助力自动驾驶鲁棒性研究
我们提出CROWD(City Road Observations With Dashcams),一个手动筛选的高质量城市行车记录视频数据集,包含51,753个片段,总计20,275.56小时(42,032个视频),覆盖全球238个国家和地区、7,103个有人居住地,横跨六大洲。数据聚焦常规驾驶行为,排除事故及编辑内容。每个片段配有时间(白天/夜晚)和车辆类型的人工标注。为降低基准测试门槛,我们提供基于YOLOv11x生成的80类MS-COCO检测结果和局部多目标跟踪(BoT-SORT),涵盖行人、自行车、汽车、交通灯等常见目标。数据以视频标识符与片段边界形式发布,不重传原始视频,保障可复现性。
原文摘要 · Abstract (English)
We introduce CROWD (City Road Observations With Dashcams), a manually curated dataset of ordinary, minute scale, temporally contiguous, unedited, front facing urban dashcam segments screened and segmented from publicly available YouTube videos. CROWD is designed to support cross-domain robustness and interaction analysis by prioritising routine driving and explicitly excluding crashes, crash aftermath, and other edited or incident-focused content. The release contains 51,753 segment records spanning 20,275.56 hours (42,032 videos), covering 7,103 named inhabited places in 238 countries and territories across all six inhabited continents (Africa, Asia, Europe, North America, South America and Oceania), with segment level manual labels for time of day (day or night) and vehicle type. To lower the barrier for benchmarking, we provide per-segment CSV files of machine-generated detections for all 80 MS-COCO classes produced with YOLOv11x, together with segment-local multi-object tracks (BoT-SORT); e.g. person, bicycle, motorcycle, car, bus, truck, traffic light, stop sign, etc. CROWD is distributed as video identifiers with segment boundaries and derived annotations, enabling reproducible research without redistributing the underlying videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。