arXiv:2501.06808cs.CV2025-01被引 22

用视觉语言模型提升遥感图像语义变化检测泛化能力

Semantic-CD: Remote Sensing Image Semantic Change Detection towards Open-vocabulary Setting

  • 引入CLIP开放词汇语义,实现跨类别泛化
  • 双时相特征+解码器分离设计,精度优于传统方法
  • 适合需快速适配新类别的遥感变化分析场景

遥感图像语义变化检测旨在识别同一地点不同时期图像中的变化区域并分类。传统方法在实际应用中难以跨语义类别泛化。为此,本文提出面向开放词汇设置的Semantic-CD方法,利用视觉语言基础模型CLIP的广泛词汇知识,增强模型对新类别的适应能力。该方法采用完全解耦的多任务学习框架,同时完成二值变化检测与语义变化检测。模型包含四个组件:双时相CLIP视觉编码器、开放语义提示器(生成开放词汇语义代价图)、二值变化检测解码器和语义变化检测解码器。在SECOND数据集上的实验表明,Semantic-CD生成更精确的掩码,显著降低语义分类错误,验证了将视觉语言先验应用于语义变化检测任务的有效性。

原文摘要 · Abstract (English)

Remote sensing image semantic change detection is a method used to analyze remote sensing images, aiming to identify areas of change as well as categorize these changes within images of the same location taken at different times. Traditional change detection methods often face challenges in generalizing across semantic categories in practical scenarios. To address this issue, we introduce a novel approach called Semantic-CD, specifically designed for semantic change detection in remote sensing images. This method incorporates the open vocabulary semantics from the vision-language foundation model, CLIP. By utilizing CLIP's extensive vocabulary knowledge, our model enhances its ability to generalize across categories and improves segmentation through fully decoupled multi-task learning, which includes both binary change detection and semantic change detection tasks. Semantic-CD consists of four main components: a bi-temporal CLIP visual encoder for extracting features from bi-temporal images, an open semantic prompter for creating semantic cost volume maps with open vocabulary, a binary change detection decoder for generating binary change detection masks, and a semantic change detection decoder for producing semantic labels. Experimental results on the SECOND dataset demonstrate that Semantic-CD achieves more accurate masks and reduces semantic classification errors, illustrating its effectiveness in applying semantic priors from vision-language foundation models to SCD tasks.

遥感变化检测视觉语言模型开放词汇多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。