arXiv:2505.17665cs.CVcs.AI2025-05

用区域注意力提升遥感图像多类别分割精度

EMRA-proxy: Enhancing Multi-Class Region Semantic Segmentation in Remote Sensing Images with Attention Proxy

  • 基于区域级注意力机制,捕捉长距离上下文关系
  • 在三个公开数据集上达到领先分割准确率
  • 适合遥感图像精细分类与复杂场景分析

高分辨率遥感(HRRS)图像分割因空间布局复杂、目标外观多样而具有挑战性。虽然卷积神经网络擅长捕捉局部特征,但难以建模长程依赖;而变换器虽能建模全局上下文,却常忽略局部细节且计算开销大。我们提出一种新方法——区域感知代理网络(RAPNet),包含两个模块:上下文区域注意力(CRA)和全局类别精炼(GCR)。不同于传统基于网格的方法,RAPNet在区域层面操作,实现更灵活的分割。CRA模块利用变换器捕捉区域级上下文依赖,生成语义区域掩码(SRM)。GCR模块学习全局类别注意力图以精炼多类信息,结合SRM与注意力图实现精确分割。在三个公开数据集上的实验表明,RAPNet优于现有最先进方法,显著提升多类别分割准确率。

原文摘要 · Abstract (English)

High-resolution remote sensing (HRRS) image segmentation is challenging due to complex spatial layouts and diverse object appearances. While CNNs excel at capturing local features, they struggle with long-range dependencies, whereas Transformers can model global context but often neglect local details and are computationally expensive.We propose a novel approach, Region-Aware Proxy Network (RAPNet), which consists of two components: Contextual Region Attention (CRA) and Global Class Refinement (GCR). Unlike traditional methods that rely on grid-based layouts, RAPNet operates at the region level for more flexible segmentation. The CRA module uses a Transformer to capture region-level contextual dependencies, generating a Semantic Region Mask (SRM). The GCR module learns a global class attention map to refine multi-class information, combining the SRM and attention map for accurate segmentation.Experiments on three public datasets show that RAPNet outperforms state-of-the-art methods, achieving superior multi-class segmentation accuracy.

遥感图像语义分割注意力机制Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。