用文本引导对比损失,让遥感图像变化描述更精准。
DFM: Difference Feature Modeling with Text-Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning

- 引入文本引导的门控对比损失,聚焦关键差异特征。
- 在多个数据集上优于现有方法,显著提升变化描述准确率。
- 适合遥感变化检测与自动描述任务的研究者参考。
遥感图像变化描述(RSICC)旨在自动生成不同时相遥感图像间变化的描述。现有模型仍依赖单一自回归生成范式,易忽略图像间具有区分性的差异特征。为此,本文重新设计训练范式,提出一种新的差异特征建模框架(DFM)。具体地,引入文本引导的门控对比损失(TGCL),从文本模态视角指导视觉编码器提取关键特征;同时结合预训练的变化检测模型,迁移稳定的检测知识。为进一步增强表征能力,设计联合特征建模(JFM)模块,融合多尺度差异表示,从而捕捉多时相图像间的完整时空变化。在多个数据集上的大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time points. Existing models still rely on a single autoregressive generation paradigm, which tends to prioritize learning easily generated vocabulary over capturing discriminative differences between images. To address this, we reframe the training paradigm and propose a novel Difference Feature Modeling (DFM) framework. Specifically, we introduce a Text-guided Gated Contrastive Loss (TGCL) to guide the vision encoder to extract critical features from a text-modal perspective. Additionally, we incorporate a pre-trained Change Detection model to transfer stable change detection knowledge. In order to further enhance the representation, we design a Joint Feature Modeling (JFM) module to achieve the fusion of multi-scale difference representations, thereby capturing comprehensive spatiotemporal variations between multi-temporal images. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。