arXiv:2509.00598cs.CV2025-09被引 1

无需训练,统一解决遥感图像语义分割难题

DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation

  • 拆分视觉与文本表示,分层对齐局部语义与全局上下文
  • 在iSAID和RRSIS-D上超越现有无训练方法,实现开集分割与指代分割
  • 首次实现无需训练的遥感图像统一分割框架,适合快速部署

视觉语言模型(VLM)弥合了视觉与语言的鸿沟,实现了超越传统视觉模型的多模态理解。然而,将VLM从自然图像域迁移到遥感(RS)图像分割仍面临巨大领域差异和任务多样性挑战,尤其在开集语义分割(OVSS)和指代表达分割(RES)中。本文提出一种无训练统一框架DGL-RSIS,解耦视觉与文本表征,并在局部语义与全局上下文层面进行视觉-语言对齐。具体而言,全局-局部解耦(GLD)模块将文本分解为局部语义标记与全局上下文标记,图像则划分为类无关的掩码提案。局部视觉-文本对齐(LVTA)模块自适应地从掩码提案中提取上下文感知视觉特征,并通过知识引导的提示工程增强文本特征,实现局部视角下的开集分割。全局视觉-文本对齐(GVTA)模块采用全局增强的Grad-CAM机制捕捉上下文线索,再经掩码选择模块将像素级激活整合为掩码级分割输出,实现全局视角下的指代分割。在iSAID(OVSS)与RRSIS-D(RES)基准上的实验表明,DGL-RSIS优于现有无训练方法。消融实验证明各模块有效性。据我们所知,这是首个无需训练的遥感图像统一分割框架,有效将自然图像预训练的语义能力迁移至遥感领域。

原文摘要 · Abstract (English)

The emergence of vision language models (VLMs) bridges the gap between vision and language, enabling multimodal understanding beyond traditional visual-only deep learning models. However, transferring VLMs from the natural image domain to remote sensing (RS) segmentation remains challenging due to the large domain gap and the diversity of RS inputs across tasks, particularly in open-vocabulary semantic segmentation (OVSS) and referring expression segmentation (RES). Here, we propose a training-free unified framework, termed DGL-RSIS, which decouples visual and textual representations and performs visual-language alignment at both local semantic and global contextual levels. Specifically, a Global-Local Decoupling (GLD) module decomposes textual inputs into local semantic tokens and global contextual tokens, while image inputs are partitioned into class-agnostic mask proposals. Then, a Local Visual-Textual Alignment (LVTA) module adaptively extracts context-aware visual features from the mask proposals and enriches textual features through knowledge-guided prompt engineering, achieving OVSS from a local perspective. Furthermore, a Global Visual-Textual Alignment (GVTA) module employs a global-enhanced Grad-CAM mechanism to capture contextual cues for referring expressions, followed by a mask selection module that integrates pixel-level activations into mask-level segmentation outputs, thereby achieving RES from a global perspective. Experiments on the iSAID (OVSS) and RRSIS-D (RES) benchmarks demonstrate that DGL-RSIS outperforms existing training-free approaches. Ablation studies further validate the effectiveness of each module. To the best of our knowledge, this is the first unified training-free framework for RS image segmentation, which effectively transfers the semantic capability of VLMs trained on natural images to the RS domain without additional training.

遥感分割视觉语言模型无训练开集分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。