提出新模型解决超广域遥感图像分割难题
SFR-Net: Learning Scale-Frustum Representations for Ultra-Wide Area Remote Sensing Image Segmentation

- 用尺度锥形表示法统一建模不同尺度的地物和上下文
- 在GID和FBPS数据集上提升mIoU达4.29%和1.72%
- 适合处理大像素量与超广地理覆盖的遥感图像分割
像素数量和地理覆盖范围是遥感图像的两个关键特征。现有分割方法通常聚焦于像素量小或像素量大但地理覆盖有限的图像。本文提出一种针对超广域(UWA)遥感图像的新分割任务,其特点为像素量大且地理覆盖极广。UWA分割的核心挑战在于同时处理尺度差异显著的地物,并保持长距离上下文语义连续性。为此,我们提出尺度锥形表示网络(SFR-Net)。受不同高度拍摄遥感图像视锥的启发,构建尺度锥形表示,实现对不同尺度地物和上下文特征的统一建模。此外,设计级联跨尺度融合机制,有效整合这些表示,在增强局部语义理解的同时保障长程上下文连续性。在GID和FBPS数据集上的实验结果表明,SFR-Net性能达到当前最优,相比最强基线方法分别提升mIoU 4.29%和1.72%。此外,所提尺度锥形表示可嵌入通用分割网络,提升分割精度与收敛速度。代码将公开于https://github.com/ChuyuZhong/SFR-Net。
原文摘要 · Abstract (English)
Pixel count and geographical coverage are two key characteristics of remote sensing images. Existing remote sensing image segmentation methods typically focus on images with either a small pixel count or a large pixel count but limited geographical coverage. In this paper, we introduce a novel segmentation task targeting ultra-wide area (UWA) remote sensing images, characterized by both a large pixel count and extremely wide geographical coverage. The core challenges of UWA segmentation lie in simultaneously handling ground objects with significantly varying scales and maintaining long-range contextual semantic continuity. To address these challenges, we propose the Scale-Frustum Representation Network (SFR-Net). Inspired by the viewing frustums of remote sensing images captured from different altitudes, we construct scale-frustum representations, enabling unified modeling of ground objects and contextual features at different scales. Furthermore, we design a cascaded cross-scale fusion mechanism to effectively integrate these representations, enhancing local semantic understanding while ensuring long-range contextual continuity. Experimental results on GID and FBPS demonstrate that SFR-Net achieves state-of-the-art performance, improving mIoU by 1.72% and 4.29%, respectively, over the strongest competing methods. In addition, the proposed scale-frustum representations can be integrated into generic segmentation networks to improve both segmentation accuracy and convergence speed. The implementation code will be publicly available at https://github.com/ChuyuZhong/SFR-Net.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。