构建多分辨率多模态数据集,评估道路协同感知中激光雷达稀疏性影响
RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception

- 设计三档激光雷达分辨率,固定其他条件实现可控对比
- 包含超10万图像与2.6万点云帧,标注22万个3D边界框
- 适合研究多模态融合、低成本道路感知系统的开发者
激光雷达正被越来越多集成到交通摄像头中,以扩大覆盖范围并缓解遮挡问题。然而,单模态及相机-激光雷达融合架构在传感器配置和场景依赖感测条件下导致的激光雷达点云稀疏性变化下的表现仍不明确。本文提出RESOLVE,一个大规模真实世界基准数据集,包含多分辨率路边激光雷达与同步相机-激光雷达感知数据,用于系统评估单模态与融合架构在路边3D检测与跟踪中的性能。RESOLVE包含超过10万张图像和2.6万帧点云,共22万个人工标注的3D边界框,采集于真实城市交叉口,涵盖多样光照与天气条件,覆盖10类交通参与者。特别地,该数据集可在保持其他传感与环境因素不变的前提下,控制评估三种激光雷达分辨率水平。这使得能够在点云分布变化(由分辨率差异、感知距离及训练-推理分辨率不匹配引起)下进行公平的跨架构比较。大量基准实验结果揭示了多模态融合如何补偿激光雷达点云稀疏性,为设计成本高效的路边多模态感知系统提供线索。数据集与基准代码已公开于https://github.com/ASU-Suo-Lab/RESOLVE。
原文摘要 · Abstract (English)
LiDAR has increasingly been integrated into traffic cameras to expand coverage and mitigate occlusion in roadside cooperative perception. However, how unimodal and camera-LiDAR fusion architectures behave under variations in LiDAR point sparsity induced by sensor configurations and scene-dependent sensing conditions remains underexplored. We introduce RESOLVE, a large-scale real-world benchmark dataset featuring multi-resolution roadside LiDAR and synchronized camera-LiDAR sensing for systematic evaluation of unimodal and fusion-based architectures in roadside 3D detection and tracking. RESOLVE contains over 100k images and 26k point cloud frames with 220k manually annotated bounding boxes, captured at a real-world urban intersection across diverse lighting and weather conditions and spanning 10 classes of traffic participants. In particular, RESOLVE enables controlled evaluation across three LiDAR resolution levels while keeping all other sensing and environmental factors fixed. This allows fair cross-architecture comparisons under point cloud distribution shifts resulting from resolution variations, sensing distance, and training-inference resolution mismatches. Results from extensive benchmark experiments reveal insights into how multimodal fusion can compensate for LiDAR point sparsity, offering clues for designing cost-efficient roadside multimodal perception. The dataset and benchmark codes are available at https://github.com/ASU-Suo-Lab/RESOLVE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。