arXiv:2412.10734cs.CV2024-12TPAMI被引 28

构建首个全向高清多模态自动驾驶数据集,支持低成本传感器方案

OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving

论文配图:OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
图 1 · 摘自论文原文
  • 融合128线激光雷达、6个摄像头和6个4D成像雷达,实现全场景感知
  • 包含1501段30秒视频,超585万组同步传感器数据,标注超51.4万个3D框
  • 提供自动化稠密占据标注与基准模型,适合研究低成本感知系统

深度学习的快速发展对自动驾驶算法的数据需求日益提升。高质量数据集是开发有效数据驱动自动驾驶解决方案的关键。下一代自动驾驶数据集需具备多模态特性,涵盖先进传感器的广泛数据覆盖、精细标注及多样场景。为此,我们提出OmniHD-Scenes,一个大规模多模态数据集,提供全方位高精度数据。该数据集融合128线激光雷达、六个相机和六个4D成像雷达,实现环境全感知。数据集包含1501段约30秒长的片段,总计超过45万帧同步图像和超过585万组同步传感器数据点。我们提出一种新型4D标注流程,已对200段片段完成标注,生成超过51.4万个精确3D边界框,并包含静态场景语义分割标注。此外,我们引入一种新型自动化密集占据真值生成管道,有效利用非关键帧信息。同时,我们建立了全面的评估指标、基线模型和基准测试,采用环视相机与4D成像雷达探索成本效益高的自动驾驶传感方案。大量实验验证了该低成本配置的有效性及其在恶劣条件下的鲁棒性。数据将公开发布于https://www.2077ai.com/OmniHD-Scenes。

原文摘要 · Abstract (English)

The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions. Next-generation autonomous driving datasets must be multimodal, incorporating data from advanced sensors that feature extensive data coverage, detailed annotations, and diverse scene representation. To address this need, we present OmniHD-Scenes, a large-scale multimodal dataset that provides comprehensive omnidirectional high-definition data. The OmniHD-Scenes dataset combines data from 128-beam LiDAR, six cameras, and six 4D imaging radar systems to achieve full environmental perception. The dataset comprises 1501 clips, each approximately 30-s long, totaling more than 450K synchronized frames and more than 5.85 million synchronized sensor data points. We also propose a novel 4D annotation pipeline. To date, we have annotated 200 clips with more than 514K precise 3D bounding boxes. These clips also include semantic segmentation annotations for static scene elements. Additionally, we introduce a novel automated pipeline for generation of the dense occupancy ground truth, which effectively leverages information from non-key frames. Alongside the proposed dataset, we establish comprehensive evaluation metrics, baseline models, and benchmarks for 3D detection and semantic occupancy prediction. These benchmarks utilize surround-view cameras and 4D imaging radar to explore cost-effective sensor solutions for autonomous driving applications. Extensive experiments demonstrate the effectiveness of our low-cost sensor configuration and its robustness under adverse conditions. Data will be released at https://www.2077ai.com/OmniHD-Scenes.

自动驾驶多模态数据4D雷达占据预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。