面向自动驾驶的低成本大尺度深度数据集,覆盖多场景多时域。
ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
- 用轻量化设备采集20万帧高分辨率图像,支持大规模部署
- 涵盖昼夜与恶劣天气,覆盖北美、欧洲、亚洲多类道路场景
- 提供完整校准与隐私处理流程,适合研究者复现和模型训练
深度估计是自动驾驶等无人系统在开放城市环境中实现空间感知的基础。现有数据集如KITTI、nuScenes和DDAD虽推动了领域发展,但在多样性、可扩展性方面受限,且性能基准已接近饱和。一个未被充分讨论的限制是传感器经济性:这些数据集依赖昂贵、耗电且难以规模化复制的多激光雷达配置,制约了地理与时间维度的多样性。本文提出ROVR,一个大规模、多样化且成本可控的深度数据集,旨在捕捉真实驾驶环境的复杂性。ROVR包含20万帧高分辨率图像,覆盖高速公路、乡村及城市场景,涵盖昼夜周期与恶劣天气条件,数据采集范围遍及北美、欧洲和亚洲。我们还公开了标定、同步、预处理及隐私处理流程,支持第三方复现。轻量级采集方案实现可扩展数据收集,稀疏但统计充分的真实标签经密度消融验证,可支撑鲁棒模型训练。大量消融实验揭示了三类共性失效模式——光照坍缩、几何混淆与距离饱和,当前架构普遍存在。数据集、数据加载器、标定与隐私处理流程、评估代码均已开源,访问地址:https://xiandaguo.net/ROVR-Open-Dataset。
原文摘要 · Abstract (English)
Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environments. Existing depth datasets such as KITTI, nuScenes, and DDAD have advanced the field but are limited in diversity and scalability, and benchmark performance on them is approaching saturation. A less discussed constraint is \emph{sensor economics}: the bespoke multi-LiDAR rigs behind these datasets are expensive, power-hungry, and difficult to replicate at fleet scale, which caps the geographic and temporal diversity that any single benchmark can cover. We present ROVR, a large-scale, diverse, and cost-efficient depth dataset designed to capture the complexity of real-world driving. ROVR comprises 200K high-resolution frames across highway, rural, and urban scenarios, spanning day/night cycles and adverse weather conditions, collected across North America, Europe, and Asia. We additionally release the calibration, synchronization, preprocessing, and privacy pipeline so that the platform can be reproduced by third parties. The lightweight acquisition pipeline enables scalable collection, while sparse but statistically sufficient ground truth -- validated by a density ablation -- supports robust model training. Extensive ablation studies further characterize performance across scene types, illumination, weather conditions, and ground-truth sparsity levels, and identify three qualitatively distinct failure modes -- photometric collapse, geometric confusion, and range saturation -- that current architectures share. The dataset, data loaders, calibration and privacy pipelines, and evaluation code are publicly available at \url{https://xiandaguo.net/ROVR-Open-Dataset}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。