首个跨多个路口的车路协同感知数据集,支持多车多基础设施协同感知。
UrbanIng-V2X: A Large-Scale Multi-Vehicle, Multi-Infrastructure Dataset Across Multiple Intersections for Cooperative Perception
- 构建三路口真实城市环境下的多车多传感器协同感知数据集。
- 包含712,000个3D边界框标注,覆盖13类目标,每序列20秒、10Hz标注。
- 适合研究车路协同、多智能体感知与真实交通场景算法验证。
近期的协同感知数据集在推动智能出行应用方面发挥了关键作用,通过智能体间信息交换,缓解遮挡问题并提升场景理解能力。尽管部分现有真实数据集同时包含车-车和车-基础设施交互,但通常局限于单一路口或单辆车。缺乏涵盖多个路口、多辆联网车辆与基础设施传感器的综合性感知数据集,限制了算法在多样化交通环境中的基准测试,导致模型可能因路口布局与交通行为相似而过拟合,表现出误导性高精度。为此,我们提出UrbanIng-V2X,首个大规模、多模态数据集,支持在德国英戈尔施塔特三个城市路口部署的车辆与基础设施传感器进行协同感知。该数据集包含34个时间对齐且空间校准的传感器序列,每段持续20秒。所有序列记录了三个路口之一的场景,涉及两辆车辆及最多三根安装于基础设施上的传感器杆,在协同场景下运行。总计提供12个车载RGB相机、2个车载激光雷达、17个基础设施热成像相机和12个基础设施激光雷达的数据。所有序列以10Hz频率标注3D边界框,覆盖13类物体,全数据集共约712,000个标注实例。我们使用先进协同感知方法进行了全面评估,并公开发布代码库、数据集、高清地图及完整数据采集环境的数字孪生。
原文摘要 · Abstract (English)
Recent cooperative perception datasets have played a crucial role in advancing smart mobility applications by enabling information exchange between intelligent agents, helping to overcome challenges such as occlusions and improving overall scene understanding. While some existing real-world datasets incorporate both vehicle-to-vehicle and vehicle-to-infrastructure interactions, they are typically limited to a single intersection or a single vehicle. A comprehensive perception dataset featuring multiple connected vehicles and infrastructure sensors across several intersections remains unavailable, limiting the benchmarking of algorithms in diverse traffic environments. Consequently, overfitting can occur, and models may demonstrate misleadingly high performance due to similar intersection layouts and traffic participant behavior. To address this gap, we introduce UrbanIng-V2X, the first large-scale, multi-modal dataset supporting cooperative perception involving vehicles and infrastructure sensors deployed across three urban intersections in Ingolstadt, Germany. UrbanIng-V2X consists of 34 temporally aligned and spatially calibrated sensor sequences, each lasting 20 seconds. All sequences contain recordings from one of three intersections, involving two vehicles and up to three infrastructure-mounted sensor poles operating in coordinated scenarios. In total, UrbanIng-V2X provides data from 12 vehicle-mounted RGB cameras, 2 vehicle LiDARs, 17 infrastructure thermal cameras, and 12 infrastructure LiDARs. All sequences are annotated at a frequency of 10 Hz with 3D bounding boxes spanning 13 object classes, resulting in approximately 712k annotated instances across the dataset. We provide comprehensive evaluations using state-of-the-art cooperative perception methods and publicly release the codebase, dataset, HD map, and a digital twin of the complete data collection environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。