针对高分辨率遥感图像分割,提出自适应分层特征精炼框架,提升细节与语义平衡。
Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

- 分阶段自适应融合,根据局部特征差异动态分配权重
- 引入频率残差模块,保留预训练特征并注入频域信息
- 结合边界、物体性与类别关系三重先验,减少混淆
高分辨率遥感图像语义分割受益于强大的预训练分层编码器,但如何有效利用其多阶段表征仍具挑战。邻近区域对细节与语义上下文的需求不同,任务特异性变换易破坏有用特征,传统语义监督提供的结构引导有限。本文提出HAFR-Net,一种渐进式精炼框架,不采用单一解码器变换替换分层表示,而是自适应组织并保守精炼。异质性引导的阶段自适应融合(HG-SAF)基于局部特征变化预测密集阶段权重;频率残差适配器(FRA)通过有界、零初始化残差分支注入频率信息,保持融合表示作为参考;混淆感知三先验解码器(CATP)则利用边界、物体性和训练生成的类别关系线索正则化预测。在匹配的Swin-B训练和单尺度推理协议下,HAFR-Net在ISPRS Vaihingen、Potsdam、LoveDA和OpenEarthMap上分别达到84.12%、87.86%、55.17%和67.70%的mIoU,较匹配的UPerNet基线分别提升0.55、0.95、1.55和1.84个百分点。控制分析进一步表明:空间重加权不仅依赖内容,且在边界与细结构精度上优于匹配的空间与光谱方法,并显著降低预定义类别对间的混淆。
原文摘要 · Abstract (English)
Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task-specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR-Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity-Guided Stage-Adaptive Fusion (HG-SAF) predicts dense stage weights conditioned on local feature variation. A Frequency-Residual Adapter (FRA) then injects frequency information through a bounded, zero-initialized residual branch that keeps the fused representation as its reference. A Confusion-Aware Tri-Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training-derived class-relation cues. Under a matched Swin-B training and single-scale inference protocol, HAFR-Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。