用文字提示引导模型精准定位变化区域,提升遥感变化检测准确率
LG-CD: Enhancing Language-Guided Change Detection through SAM2 Adaptation
- 基于SAM2构建多尺度特征提取框架,融合文本与视觉信息
- 在三个数据集上实现最优性能,最高达98.6%的F1分数
- 适合需要高精度变化检测的遥感应用开发者
遥感变化检测(RSCD)通常通过分析多时相图像来识别地表覆盖或状态的变化。当前大多数深度学习方法主要关注单一视觉信息,忽视了文本等多模态数据提供的丰富语义信息。为此,我们提出一种新型语言引导变化检测模型(LG-CD)。该模型利用自然语言提示引导网络关注感兴趣区域,显著提升检测的准确性和鲁棒性。具体而言,LG-CD采用视觉基础模型SAM2作为特征提取器,从双时相遥感图像中捕获从高分辨率到低分辨率的多尺度金字塔特征。随后,通过多层适配器对模型进行微调,以适应下游任务。设计了文本融合注意力模块(TFAM),用于对齐视觉与文本信息,使模型能借助文本提示聚焦目标变化区域。最后,引入视觉-语义融合解码器(V-SFD),通过交叉注意力机制深度融合视觉与语义信息,生成高精度变化检测掩码。在LEVIR-CD、WHU-CD和SYSU-CD三个数据集上的实验表明,LG-CD持续优于现有最先进方法。此外,该方法为利用多模态信息实现泛化变化检测提供了新思路。
原文摘要 · Abstract (English)
Remote Sensing Change Detection (RSCD) typically identifies changes in land cover or surface conditions by analyzing multi-temporal images. Currently, most deep learning-based methods primarily focus on learning unimodal visual information, while neglecting the rich semantic information provided by multimodal data such as text. To address this limitation, we propose a novel Language-Guided Change Detection model (LG-CD). This model leverages natural language prompts to direct the network's attention to regions of interest, significantly improving the accuracy and robustness of change detection. Specifically, LG-CD utilizes a visual foundational model (SAM2) as a feature extractor to capture multi-scale pyramid features from high-resolution to low-resolution across bi-temporal remote sensing images. Subsequently, multi-layer adapters are employed to fine-tune the model for downstream tasks, ensuring its effectiveness in remote sensing change detection. Additionally, we design a Text Fusion Attention Module (TFAM) to align visual and textual information, enabling the model to focus on target change regions using text prompts. Finally, a Vision-Semantic Fusion Decoder (V-SFD) is implemented, which deeply integrates visual and semantic information through a cross-attention mechanism to produce highly accurate change detection masks. Our experiments on three datasets (LEVIR-CD, WHU-CD, and SYSU-CD) demonstrate that LG-CD consistently outperforms state-of-the-art change detection methods. Furthermore, our approach provides new insights into achieving generalized change detection by leveraging multimodal information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。