arXiv:2604.02010cs.CV2026-04

提出DR-Seg框架,分离语义与结构特征,精准增强遥感分割边界。

Decouple and Rectify: Semantics-Preserving Structural Enhancement for Open-Vocabulary Remote Sensing Segmentation

  • 将CLIP特征分解为语义与结构子空间,实现定向增强。
  • 在8个基准上达到新最优,边界分割精度显著提升。
  • 适合需要高精度遥感图像语义分割的研究者使用。

遥感领域开放词汇语义分割需兼顾语言对齐识别与细粒度空间划分。尽管CLIP具备强大语义泛化能力,但其全局对齐的视觉表示难以捕捉结构细节。现有方法尝试引入预训练于遥感数据的DINO特征进行补偿,但将CLIP视为整体语义空间,无法定位所需结构增强区域,易破坏语义一致性且难以精确划分边界。本文提出DR-Seg——一种解耦与修正框架。核心观察为:CLIP特征通道存在功能异质性而非均匀语义空间。基于此,DR-Seg将特征解耦为以语义为主和以结构为主的子空间,使DINO可针对性增强结构信息而不干扰语义。随后,先验驱动的图修正模块在DINO引导下注入高保真结构先验,生成优化分支;不确定性引导的自适应融合模块动态整合该分支与原始CLIP分支,完成最终预测。在八个基准上的全面实验表明,DR-Seg达到当前最优性能。

原文摘要 · Abstract (English)

Open-vocabulary semantic segmentation in the remote sensing (RS) field requires both language-aligned recognition and fine-grained spatial delineation. Although CLIP offers robust semantic generalization, its global-aligned visual representations inherently struggle to capture structural details. Recent methods attempt to compensate for this by introducing RS-pretrained DINO features. However, these methods treat CLIP representations as a monolithic semantic space and cannot localize where structural enhancement is required, failing to effectively delineate boundaries while risking the disruption of CLIP's semantic integrity. To address this limitation, we propose DR-Seg, a novel decouple-and-rectify framework in this paper. Our method is motivated by the key observation that CLIP feature channels exhibit distinct functional heterogeneity rather than forming a uniform semantic space. Building on this insight, DR-Seg decouples CLIP features into semantics-dominated and structure-dominated subspaces, enabling targeted structural enhancement by DINO without distorting language-aligned semantics. Subsequently, a prior-driven graph rectification module injects high-fidelity structural priors under DINO guidance to form a refined branch, while an uncertainty-guided adaptive fusion module dynamically integrates this refined branch with the original CLIP branch for final prediction. Comprehensive experiments across eight benchmarks demonstrate that DR-Seg establishes a new state-of-the-art.

遥感分割语义分割CLIP结构增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。