arXiv:2509.03961cs.CVcs.AI2025-09被引 9

融合图像与文本信息,提升遥感变化检测精度与抗干扰能力

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

  • 引入图文双模态,通过文本增强语义理解能力
  • 在三个数据集上均超越现有方法,最高提升3.8%的F1分数
  • 适合需要高鲁棒性的遥感变化检测场景

尽管深度学习已推动遥感变化检测(RSCD)发展,但多数方法仅依赖图像模态,限制了特征表达、变化模式建模和泛化能力,尤其在光照与噪声干扰下表现不佳。为此,我们提出MMChange,一种结合图像与文本模态的多模态RSCD方法。设计图像特征优化(IFR)模块以突出关键区域并抑制环境噪声;为克服图像特征的语义局限,利用视觉语言模型(VLM)生成双时相图像的语义描述;进一步通过文本差异增强(TDE)模块捕捉细粒度语义变化,引导模型关注有意义的变化。为弥合模态异构性,构建图像-文本特征融合(ITFF)模块实现深层跨模态融合。在LEVIRCD、WHUCD和SYSUCD数据集上的大量实验表明,MMChange在多个指标上持续优于现有先进方法,验证了其在多模态变化检测中的有效性。代码已开源:https://github.com/yikuizhai/MMChange。

原文摘要 · Abstract (English)

Although deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization especially under illumination and noise disturbances. To address this, we propose MMChange, a multimodal RSCD method that combines image and text modalities to enhance accuracy and robustness. An Image Feature Refinement (IFR) module is introduced to highlight key regions and suppress environmental noise. To overcome the semantic limitations of image features, we employ a vision language model (VLM) to generate semantic descriptions of bitemporal images. A Textual Difference Enhancement (TDE) module then captures fine grained semantic shifts, guiding the model toward meaningful changes. To bridge the heterogeneity between modalities, we design an Image Text Feature Fusion (ITFF) module that enables deep cross modal integration. Extensive experiments on LEVIRCD, WHUCD, and SYSUCD demonstrate that MMChange consistently surpasses state of the art methods across multiple metrics, validating its effectiveness for multimodal RSCD. Code is available at: https://github.com/yikuizhai/MMChange.

遥感变化检测多模态融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。