arXiv:2603.17920cs.CV2026-03被引 2

用3D几何自动标注无人机航拍图像,大幅降低人工成本。

SegFly: A Dataset and 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale

  • 通过2D→3D→2D流程,从少量人工标注图生成海量伪标签。
  • 自动生成97%的RGB标签和100%的热成像标签,准确率达88%以上。
  • 适合做多模态无人机场景理解的研究者,尤其关注数据标注效率。

无人机语义分割对空中场景理解至关重要,但现有RGB与RGB-T数据集受限于规模、多样性及标注效率,主要因人工标注成本高,且普通无人机难以实现精准的RGB-T配准。为此,我们提出一种可扩展的几何驱动2D-3D-2D范式,利用高重叠航拍图像的多视角冗余,将少量手动标注的RGB图像提升至语义3D点云,并渲染到所有视角,实现统一框架下对RGB与热成像模态的自动标签传播。仅需标注不足3%的RGB图像,即可生成覆盖大规模图像集的密集伪真值,自动产出97%的RGB标签与100%的热成像标签,且无需2D人工修正即达91%与88%的标注准确率。进一步将该范式拓展至跨模态图像配准,以3D几何为中间对齐空间,实现完全自动、强像素级的RGB-T对齐,准确率达87%,无需硬件同步。基于此框架,我们在已有地理参考航拍影像上构建了SegFly数据集,包含超2万张高分辨率RGB图像及超过1.5万对几何对齐的RGB-T图像,覆盖城市、工业与农村多种环境,跨越多个高度与季节。在该数据集上,我们建立Firefly基线模型,验证传统架构与视觉基础模型均显著受益于SegFly监督,凸显几何驱动2D-3D-2D流水线在可扩展多模态空中场景理解中的潜力。数据与代码见https://github.com/markus-42/SegFly。

原文摘要 · Abstract (English)

Semantic segmentation for uncrewed aerial vehicles (UAVs) is fundamental for aerial scene understanding, yet existing RGB and RGB-T datasets remain limited in scale, diversity, and annotation efficiency due to the high cost of manual labeling and the difficulties of accurate RGB-T alignment on off-the-shelf UAVs. To address these challenges, we propose a scalable geometry-driven 2D-3D-2D paradigm that leverages multi-view redundancy in high-overlap aerial imagery to automatically propagate labels from a small subset of manually annotated RGB images to both RGB and thermal modalities within a unified framework. By lifting less than 3% of RGB images into a semantic 3D point cloud and rendering it into all views, our approach enables dense pseudo ground-truth generation across large image collections, automatically producing 97% of RGB labels and 100% of thermal labels while achieving 91% and 88% annotation accuracy without any 2D manual refinement. We further extend this 2D-3D-2D paradigm to cross-modal image registration, using 3D geometry as an intermediate alignment space to obtain fully automatic, strong pixel-level RGB-T alignment with 87% registration accuracy and no hardware-level synchronization. Applying our framework to existing geo-referenced aerial imagery, we construct SegFly, a large-scale benchmark with over 20,000 high-resolution RGB images and more than 15,000 geometrically aligned RGB-T pairs spanning diverse urban, industrial, and rural environments across multiple altitudes and seasons. On SegFly, we establish the Firefly baseline for RGB and thermal semantic segmentation and show that both conventional architectures and vision foundation models benefit substantially from SegFly supervision, highlighting the potential of geometry-driven 2D-3D-2D pipelines for scalable multi-modal aerial scene understanding. Data and Code available at https://github.com/markus-42/SegFly.

无人机多模态语义分割数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。