首个支持2D/3D多模态的无人机场景理解数据集,推动高阶感知任务发展。
UAVScenes: A Multi-Modal Dataset for UAVs
- 基于MARS-LVIG扩展,提供图像与点云逐帧语义标注
- 包含6-DoF位姿与精细标注,支持分割、深度估计等10+任务
- 适合研究无人机多模态感知、场景理解的科研人员使用
多模态感知对无人机运行至关重要,能全面理解周围环境。然而现有大多数多模态无人机数据集主要偏向定位与三维重建任务,或仅支持地图级语义分割,因缺乏图像与激光雷达点云的逐帧标注,难以支持高层场景理解任务。为填补这一空白并推动多模态无人机感知发展,我们提出UAVScenes,一个大规模数据集,用于跨2D与3D模态的多种任务基准测试。该数据集基于原始仅用于同时定位与地图构建(SLAM)的高质量多模态无人机数据集MARS-LVIG进行增强,新增逐帧图像与激光雷达点云的语义标注,并提供精确的6-度自由度(6-DoF)位姿。这些改进使数据集可支持分割、深度估计、6-DoF定位、场景识别及新视角合成(NVS)等多种任务。数据集已开源:https://github.com/sijieaaa/UAVScenes。
原文摘要 · Abstract (English)
Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward localization and 3D reconstruction tasks, or only support map-level semantic segmentation due to the lack of frame-wise annotations for both camera images and LiDAR point clouds. This limitation prevents them from being used for high-level scene understanding tasks. To address this gap and advance multi-modal UAV perception, we introduce UAVScenes, a large-scale dataset designed to benchmark various tasks across both 2D and 3D modalities. Our benchmark dataset is built upon the well-calibrated multi-modal UAV dataset MARS-LVIG, originally developed only for simultaneous localization and mapping (SLAM). We enhance this dataset by providing manually labeled semantic annotations for both frame-wise images and LiDAR point clouds, along with accurate 6-degree-of-freedom (6-DoF) poses. These additions enable a wide range of UAV perception tasks, including segmentation, depth estimation, 6-DoF localization, place recognition, and novel view synthesis (NVS). Our dataset is available at https://github.com/sijieaaa/UAVScenes
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。