arXiv:2602.23869cs.CV2026-02

无需训练即可实现遥感图像开放词汇语义分割

Open-Vocabulary Semantic Segmentation in Remote Sensing via Hierarchical Attention Masking and Model Composition

  • 用SAM生成的分层掩码约束CLIP自注意力,提升多尺度感知
  • 通过组合多个遥感专用CLIP模型,跨基准达到领先性能
  • 适合遥感图像分割研究者,尤其关注零样本场景应用

本文提出ReSeg-CLIP,一种无需训练的遥感开放词汇语义分割方法。为解决视觉语言模型(如CLIP)在语义分割中因自注意力层内交互不当导致的问题,我们引入基于SAM生成掩码的分层约束机制,在多尺度上限制不必要交互。同时提出模型组合策略,通过新权重方案评估不同文本提示下的表征质量,平均多个遥感专用CLIP变体的参数。该方法在三个遥感基准数据集上均取得当前最优性能,且无需额外训练。

原文摘要 · Abstract (English)

In this paper, we propose ReSeg-CLIP, a new training-free Open-Vocabulary Semantic Segmentation method for remote sensing data. To compensate for the problems of vision language models, such as CLIP in semantic segmentation caused by inappropriate interactions within the self-attention layers, we introduce a hierarchical scheme utilizing masks generated by SAM to constrain the interactions at multiple scales. We also present a model composition approach that averages the parameters of multiple RS-specific CLIP variants, taking advantage of a new weighting scheme that evaluates representational quality using varying text prompts. Our method achieves state-of-the-art results across three RS benchmarks without additional training.

遥感分割开放词汇CLIP零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。