arXiv:2506.21866cs.CV2025-06IJCAI被引 9

融合卷积与Transformer优势,提升遥感图像目标分割精度

Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing Images

  • 双视角注意力机制捕捉全局与局部特征,傅里叶空间融合减少偏差
  • 在WHU-RS12、SpaceNet、ISPRS等数据集上达到最高分割精度
  • 适合遥感图像分析、复杂场景分割任务的研究者与工程师使用

自动分割光学遥感图像(ORSIs)中的目标是一项重要任务。现有模型多基于卷积或Transformer特征,各有优势,但同时利用两者面临特征异质性、模型复杂度高和参数量大等挑战,且常被忽视,导致分割性能不佳。为此,我们提出一种新型双视角统一Transformer(DPU-Former),其独特结构可同步整合长程依赖与空间细节。设计了全局-局部混合注意力机制,通过双重视角捕获多样信息,并引入傅里叶空间融合策略,实现高效特征融合。此外,采用门控线性前馈网络增强表达能力。还构建了DPU-Former解码器,用于聚合与强化多层特征。实验表明,该模型在多个数据集上超越现有最先进方法。代码已开源:https://github.com/CSYSI/DPU-Former。

原文摘要 · Abstract (English)

Automatically segmenting objects from optical remote sensing images (ORSIs) is an important task. Most existing models are primarily based on either convolutional or Transformer features, each offering distinct advantages. Exploiting both advantages is valuable research, but it presents several challenges, including the heterogeneity between the two types of features, high complexity, and large parameters of the model. However, these issues are often overlooked in existing the ORSIs methods, causing sub-optimal segmentation. For that, we propose a novel Dual-Perspective United Transformer (DPU-Former) with a unique structure designed to simultaneously integrate long-range dependencies and spatial details. In particular, we design the global-local mixed attention, which captures diverse information through two perspectives and introduces a Fourier-space merging strategy to obviate deviations for efficient fusion. Furthermore, we present a gated linear feed-forward network to increase the expressive ability. Additionally, we construct a DPU-Former decoder to aggregate and strength features at different layers. Consequently, the DPU-Former model outperforms the state-of-the-art methods on multiple datasets. Code: https://github.com/CSYSI/DPU-Former.

遥感图像目标分割Transformer双视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。