arXiv:2604.26893cs.CV2026-04

解决无人机红外可见光图像错位下的精细分割难题

Graph-based Semantic Calibration Network for Unaligned UAV RGBT Image Semantic Segmentation and A Large-scale Benchmark

论文配图:Graph-based Semantic Calibration Network for Unaligned UAV RGBT Image Semantic Segmentation and A Large-scale Benchmark
图 1 · 摘自论文原文
  • 用图结构建模类别层级与共现关系,提升相似类别的区分能力
  • 通过共享空间变形对齐,减少模态干扰,改善错位问题
  • 构建超2.5万张图像的大型基准数据集,覆盖61种地物类别

细粒度的RGBT图像语义分割对全天候无人机场景理解至关重要。然而,无人机RGBT图像分割面临两个耦合挑战:传感器视差和平台振动导致的跨模态空间错位,以及俯视视角下细粒度地物间的严重语义混淆。为此,本文提出面向未对齐无人机RGBT图像分割的图结构语义校准网络(GSCNet)。设计特征解耦与对齐模块(FDAM),将各模态分解为共享结构与私有感知成分,在共享子空间中进行可变形对齐,实现鲁棒的空间校正并降低模态外观干扰。提出语义图校准模块(SGCM),将无人机场景中地物类别的层次化分类体系与共现规律编码为结构化类别图,并通过图注意力机制融入推理过程,校准视觉相似及稀有类别的预测。此外,构建了目前最大且最细粒度的未对齐无人机多模态分割基准数据集URTF,包含超过25,000对图像、61个语义类别,具有真实跨模态错位。在URTF上的大量实验表明,GSCNet显著优于现有方法,尤其在细粒度类别上表现突出。数据集开源地址:https://github.com/mmic-lcl/Datasets-and-benchmark-code。

原文摘要 · Abstract (English)

Fine-grained RGBT image semantic segmentation is crucial for all-weather unmanned aerial vehicle (UAV) scene understanding. However, UAV RGBT image semantic segmentation faces two coupled challenges: cross-modal spatial misalignment caused by sensor parallax and platform vibration, and severe semantic confusion among fine-grained ground objects under top-down aerial views. To address these issues, we propose a Graph-based Semantic Calibration Network (GSCNet) for unaligned UAV RGBT image semantic segmentation. Specifically, we design a Feature Decoupling and Alignment Module (FDAM) that decouples each modality into shared structural and private perceptual components and performs deformable alignment in the shared subspace, enabling robust spatial correction with reduced modality appearance interference. Moreover, we propose a Semantic Graph Calibration Module (SGCM) that explicitly encodes the hierarchical taxonomy and co-occurrence regularities among ground-object categories in UAV scenes into a structured category graph, and incorporates these priors into graph-attention reasoning to calibrate predictions of visually similar and rare categories. In addition, we construct the Unaligned RGB-Thermal Fine-grained (URTF) benchmark, to the best of our knowledge, the largest and most fine-grained benchmark for unaligned UAV RGBT image semantic segmentation, containing over 25,000 image pairs across 61 semantic categories with realistic cross-modal misalignment. Extensive experiments on URTF demonstrate that GSCNet significantly outperforms state-of-the-art methods, with notable gains on fine-grained categories. The dataset is available at https://github.com/mmic-lcl/Datasets-and-benchmark-code.

无人机多模态语义分割图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。