arXiv:2506.21945cs.CVcs.AI2025-06

针对高分辨率遥感图像语义分割难题,提出堆叠残差网络提升分割精度。

SDRNET: Stacked Deep Residual Network for Accurate Semantic Segmentation of Fine-Resolution Remotely Sensed Images

  • 采用双编码器-解码器堆叠结构,融合长程语义与空间信息
  • 引入空洞残差块捕捉全局依赖,缓解深层网络细节丢失问题
  • 在Vaihingen和Potsdam数据集上优于现有深度网络

高分辨率遥感图像的语义分割生成的地表覆盖图在摄影测量与遥感领域备受关注。随着传感与成像技术进步,大量细粒度遥感图像(FRRS)已可获取,但其准确分割受类别差异大、关键地物遮挡导致不可见以及目标尺寸变化等因素严重影响。尽管深度卷积神经网络(DCNN)在图像特征学习方面潜力巨大,但从FRRS图像中提取足够特征以实现精确分割仍具挑战。需使模型学习鲁棒特征并生成充分的特征描述符。具体而言,需学习多上下文特征以覆盖地面场景中不同尺度目标,并利用全局-局部上下文缓解类别不平衡问题。更深层网络因逐步下采样过程导致空间细节严重丢失,进而影响分割效果与边界精度。本文提出一种堆叠深度残差网络(SDRNet),用于从FRRS图像中进行语义分割。该框架通过两个堆叠的编码器-解码器网络捕获长程语义信息的同时保留空间细节,并在每个编码器与解码器之间引入空洞残差块(DRB),以增强对全局依赖的建模能力,从而提升分割性能。在ISPRS Vaihingen和Potsdam数据集上的实验结果表明,SDRNet在语义分割任务中表现有效且具有竞争力。

原文摘要 · Abstract (English)

Land cover maps generated from semantic segmentation of high-resolution remotely sensed images have drawn mucon in the photogrammetry and remote sensing research community. Currently, massive fine-resolution remotely sensed (FRRS) images acquired by improving sensing and imaging technologies become available. However, accurate semantic segmentation of such FRRS images is greatly affected by substantial class disparities, the invisibility of key ground objects due to occlusion, and object size variation. Despite the extraordinary potential in deep convolutional neural networks (DCNNs) in image feature learning and representation, extracting sufficient features from FRRS images for accurate semantic segmentation is still challenging. These challenges demand the deep learning models to learn robust features and generate sufficient feature descriptors. Specifically, learning multi-contextual features to guarantee adequate coverage of varied object sizes from the ground scene and harnessing global-local contexts to overcome class disparities challenge even profound networks. Deeper networks significantly lose spatial details due to gradual downsampling processes resulting in poor segmentation results and coarse boundaries. This article presents a stacked deep residual network (SDRNet) for semantic segmentation from FRRS images. The proposed framework utilizes two stacked encoder-decoder networks to harness long-range semantics yet preserve spatial information and dilated residual blocks (DRB) between each encoder and decoder network to capture sufficient global dependencies thus improving segmentation performance. Our experimental results obtained using the ISPRS Vaihingen and Potsdam datasets demonstrate that the SDRNet performs effectively and competitively against current DCNNs in semantic segmentation.

语义分割遥感图像深度网络残差结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。