构建40万帧多模态无人机感知数据集,解决真实数据难获取问题
UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data
- 用高保真仿真生成跨场景、多天气的40万帧多模态数据
- 包含5种传感器数据,每帧有2D/3D框与6自由度姿态标注
- 适合做无人机检测、跟踪与轨迹预测的研究者使用
低空复杂环境下对无人机的精准感知对空域安全和智能系统至关重要。可靠解决方案的开发依赖于大规模、高精度标注的多模态数据。然而,受空域管制、隐私顾虑和环境变化影响,真实无人机数据采集存在固有局限;而3D姿态与跨模态对应关系的人工标注耗时且昂贵。为此,我们提出UAV-MM3D,一个面向低空无人机感知与运动理解的高保真多模态合成数据集。该数据集涵盖40万帧同步图像,覆盖城市、郊区、森林、沿海等多样化场景及晴朗、多云、雨天、雾天等多种天气条件,包含微型、小型、中型等多种无人机模型,支持五种模态:RGB、红外(IR)、LiDAR、雷达(Radar)与动态视觉传感器(DVS)。每帧提供2D/3D边界框、6-自由度(6-DoF)位姿与实例级标注,支持3D检测、姿态估计、目标跟踪及短时轨迹预测等核心任务。我们进一步提出LGFusionNet(LiDAR引导的多模态融合基线)与专用无人机轨迹预测基线,以促进基准测试。凭借可控的仿真环境、全面的场景覆盖与丰富标注,UAV-MM3D为推进无人机3D感知研究提供了公开基准。
原文摘要 · Abstract (English)
Accurate perception of UAVs in complex low-altitude environments is critical for airspace security and related intelligent systems. Developing reliable solutions requires large-scale, accurately annotated, and multimodal data. However, real-world UAV data collection faces inherent constraints due to airspace regulations, privacy concerns, and environmental variability, while manual annotation of 3D poses and cross-modal correspondences is time-consuming and costly. To overcome these challenges, we introduce UAV-MM3D, a high-fidelity multimodal synthetic dataset for low-altitude UAV perception and motion understanding. It comprises 400K synchronized frames across diverse scenes (urban areas, suburbs, forests, coastal regions) and weather conditions (clear, cloudy, rainy, foggy), featuring multiple UAV models (micro, small, medium-sized) and five modalities - RGB, IR, LiDAR, Radar, and DVS (Dynamic Vision Sensor). Each frame provides 2D/3D bounding boxes, 6-DoF poses, and instance-level annotations, enabling core tasks related to UAVs such as 3D detection, pose estimation, target tracking, and short-term trajectory forecasting. We further propose LGFusionNet, a LiDAR-guided multimodal fusion baseline, and a dedicated UAV trajectory prediction baseline to facilitate benchmarking. With its controllable simulation environment, comprehensive scenario coverage, and rich annotations, UAV3D offers a public benchmark for advancing 3D perception of UAVs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。