arXiv:2508.06224cs.CV2025-08中稿 · IEEE GRSL被引 2

提出TEFormer模型,提升城市遥感图像语义分割精度。

TEFormer: Texture-Aware and Edge-Guided Transformer for Semantic Segmentation of Urban Remote Sensing Images

  • 引入纹理感知模块与边缘引导解码器,增强细粒度区分能力。
  • 在Potsdam和Vaihingen数据集上分别达到88.57%和81.46%的mIoU。
  • 适用于复杂城市场景中边界模糊、纹理相似的物体分割任务。

城市遥感图像(URSIs)的精确语义分割对城市规划与环境监测至关重要。然而,由于地理对象间纹理差异细微、空间结构相似,常导致语义混淆与误分类。此外,不规则形状、模糊边界及重叠分布进一步加剧边缘形态的多样性与复杂性。为此,我们提出TEFormer:一种纹理感知且边缘引导的Transformer。其编码器包含纹理感知模块(TaM),可捕捉视觉相似类别间的细粒度纹理差异,提升语义判别力;解码器采用边缘引导三分支结构(Eg3Head),兼顾局部边缘细节与多尺度上下文信息;最后通过边缘引导特征融合模块(EgFFM)有效整合上下文、细节与边缘信息,实现精细化分割。大量实验表明,TEFormer在Potsdam数据集上达88.57% mIoU,优于次优方法0.73%;在Vaihingen上达81.46% mIoU,领先0.22%;在LoveDA数据集上以53.55%总体mIoU位列第二,仅落后最优结果0.19%。

原文摘要 · Abstract (English)

Accurate semantic segmentation of urban remote sensing images (URSIs) is essential for urban planning and environmental monitoring. However, it remains challenging due to the subtle texture differences and similar spatial structures among geospatial objects, which cause semantic ambiguity and misclassification. Additional complexities arise from irregular object shapes, blurred boundaries, and overlapping spatial distributions of objects, resulting in diverse and intricate edge morphologies. To address these issues, we propose TEFormer, a texture-aware and edge-guided Transformer. Our model features a texture-aware module (TaM) in the encoder to capture fine-grained texture distinctions between visually similar categories, thereby enhancing semantic discrimination. The decoder incorporates an edge-guided tri-branch decoder (Eg3Head) to preserve local edges and details while maintaining multiscale context-awareness. Finally, an edge-guided feature fusion module (EgFFM) effectively integrates contextual, detail, and edge information to achieve refined semantic segmentation. Extensive evaluation demonstrates that TEFormer yields mIoU scores of 88.57% on Potsdam and 81.46% on Vaihingen, exceeding the next best methods by 0.73% and 0.22%. On the LoveDA dataset, it secures the second position with an overall mIoU of 53.55%, trailing the optimal performance by a narrow margin of 0.19%.

语义分割遥感图像Transformer边缘引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。