首个基于相机的航拍语义场景补全基准,解决高空感知数据稀缺问题。
OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective
- 用相机替代激光雷达,自动将少量标注图像映射到三维点云中
- 构建包含2万余样本的航拍数据集,覆盖21类场景且跨季节与高度
- 揭示当前视觉模型在空中场景下的性能瓶颈,推动航拍3D理解发展
语义场景补全(SSC)对移动机器人三维感知至关重要,通过联合估计密集体素占据与每体素语义实现全局场景理解。尽管陆地领域(如自动驾驶)已有广泛研究,但无人机等空中应用仍处于空白状态,限制了下游任务进展。现有方法多依赖激光雷达,而其在无人机上受限于法规、重量、能耗及高空视角导致的点云稀疏性。为此,我们提出无需激光雷达的相机数据生成框架,利用经典三维重建技术,仅需<10%的标注图像即可实现语义标签向重建点云的自动迁移,大幅降低人工三维标注成本。基于此,我们推出OccuFly——首个真实世界、基于相机的航拍语义场景补全基准,覆盖城市、工业和乡村环境,采集于多个高度与全年季节。数据集包含超过20,000个样本,提供图像、语义体素网格与度量深度图,涵盖21个语义类别,遵循标准数据组织格式,便于集成。我们在OccuFly上对SSC与单目度量深度估计进行基准测试,揭示当前视觉基础模型在空中场景中的根本局限,为鲁棒空中三维理解设立新挑战。访问 https://github.com/markus-42/occufly。
原文摘要 · Abstract (English)
Semantic Scene Completion (SSC) is essential for 3D perception in mobile robotics, as it enables holistic scene understanding by jointly estimating dense volumetric occupancy and per-voxel semantics. Although SSC has been widely studied in terrestrial domains such as autonomous driving, aerial settings like autonomous flying remain largely unexplored, thereby limiting progress on downstream applications. Furthermore, LiDAR sensors are the primary modality for SSC data generation, which poses challenges for most uncrewed aerial vehicles (UAVs) due to flight regulations, mass and energy constraints, and the sparsity of LiDAR point clouds from elevated viewpoints. To address these limitations, we propose a LiDAR-free, camera-based data generation framework. By leveraging classical 3D reconstruction, our framework automates semantic label transfer by lifting <10% of annotated images into the reconstructed point cloud, substantially minimizing manual 3D annotation effort. Based on this framework, we introduce OccuFly, the first real-world, camera-based aerial SSC benchmark, captured across multiple altitudes and all seasons. OccuFly provides over 20,000 samples of images, semantic voxel grids, and metric depth maps across 21 semantic classes in urban, industrial, and rural environments, and follows established data organization for seamless integration. We benchmark both SSC and metric monocular depth estimation on OccuFly, revealing fundamental limitations of current vision foundation models in aerial settings and establishing new challenges for robust 3D scene understanding in the aerial domain. Visit https://github.com/markus-42/occufly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。