arXiv:2601.21159cs.CV2026-01

无需训练即可精准分割遥感图像中尺度多变的物体。

Spatial-Regularization-Aware Dual-Branch Collaborative Inference for Training-Free OVSS in Remote Sensing Imagery

  • 双分支协同推理,通过交叉注意力融合特征
  • 迭代扩散优化分割置信度,提升边界精度
  • 结合超像素结构,适合遥感图像细粒度分割

高分辨率遥感图像包含密集分布、尺度差异显著且边界复杂的物体,对语义分割模型的几何定位与语义预测能力提出更高要求。现有无训练开放词汇语义分割(OVSS)方法通常采用单向注入和浅层后处理策略融合对比语言-图像预训练(CLIP)与视觉基础模型(VFMs),难以满足需求。为此,我们提出一种空间正则感知的双分支协同推理框架SDCI。首先,在特征编码阶段引入交叉模型注意力融合(CAF)模块,通过相互注入自注意力图实现协同推理。其次,提出双向跨图扩散精炼(BCDR)模块,通过迭代随机游走扩散增强双分支分割得分的可靠性。最后,结合低层超像素结构,设计基于凸优化的超像素协同预测(CSCP)机制,进一步精化物体边界。在多个遥感语义分割基准上的实验表明,该方法优于现有方法。代码已公开于https://github.com/yu-ni1989/SDCI。

原文摘要 · Abstract (English)

High-resolution remote sensing images contain densely distributed objects with pronounced scale variations and complex boundaries, which impose higher demands on both the geometric localization and semantic prediction capabilities of semantic segmentation models. Existing training-free open-vocabulary semantic segmentation (OVSS) methods typically fuse Contrastive Language-Image Pretraining (CLIP) and vision foundation models (VFMs) using one-way injection and shallow post-processing strategies, making it difficult to satisfy these requirements. To address this issue, we propose a spatial-regularization-aware dual-branch collaborative inference framework for training-free OVSS, termed SDCI. First, during feature encoding, SDCI introduces a cross-model attention fusion (CAF) module, which guides collaborative inference by injecting self-attention maps into each other. Second, we propose a bidirectional cross-graph diffusion refinement (BCDR) module that enhances the reliability of dual-branch segmentation scores through iterative random-walk diffusion. Finally, we incorporate low-level superpixel structures and develop a convex-optimization-based superpixel collaborative prediction (CSCP) mechanism to further refine object boundaries. Experiments on multiple remote sensing semantic segmentation benchmarks demonstrate that our method achieves better performance than existing approaches. Our code is available at https://github.com/yu-ni1989/SDCI.

遥感分割开放词汇无训练双分支

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。