arXiv:2508.01555eess.IVcs.CV2025-08被引 4

用图文图结构增强遥感变化检测,提升语义理解能力。

MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection

  • 构建图文图条件重建机制,融合视觉与文本语义
  • 在四个公开数据集上优于主流方法,最高提升4.2%
  • 适合关注遥感图像语义分析的科研与应用人员

随着遥感卫星技术和深度学习的快速发展,遥感变化检测(RSCD)已成为区域监测的关键技术。传统方法和基于深度学习的方法虽在变化分析中取得显著进展,但在多模态数据探索与应用方面仍存在局限。为此,本文提出多模态图条件视觉-语言重构网络(MGCR-Net),进一步挖掘多模态数据的语义交互能力。多模态大语言模型(MLLM)因其出色的视觉-语言理解与对话交互能力备受关注。我们设计基于MLLM的优化策略,从原始变化检测图像生成多模态文本数据,作为MGCR的文本输入。通过双编码器框架提取视觉与文本特征。首次在RSCD任务中引入多模态图条件视觉-语言重构机制,结合图注意力构建语义图条件重构模块(SGCM),利用图结构生成视觉-语言(VL)令牌,并通过多头注意力实现视觉与文本特征的跨维度交互。重构后的VL特征经语言视觉变换器(LViT)深度融合,实现细粒度特征对齐与高层语义交互。在四个公开数据集上的实验结果表明,MGCR-Net相比主流方法表现更优,最高提升4.2%。代码已开源:https://github.com/cn-xvkong/MGCR

原文摘要 · Abstract (English)

With the advancement of remote sensing satellite technology and the rapid progress of deep learning, remote sensing change detection (RSCD) has become a key technique for regional monitoring. Traditional change detection (CD) methods and deep learning-based approaches have made significant contributions to change analysis and detection, however, many outstanding methods still face limitations in the exploration and application of multimodal data. To address this, we propose the multimodal graph-conditioned vision-language reconstruction network (MGCR-Net) to further explore the semantic interaction capabilities of multimodal data. Multimodal large language models (MLLM) have attracted widespread attention for their outstanding performance in computer vision, particularly due to their powerful visual-language understanding and dialogic interaction capabilities. Specifically, we design a MLLM-based optimization strategy to generate multimodal textual data from the original CD images, which serve as textual input to MGCR. Visual and textual features are extracted through a dual encoder framework. For the first time in the RSCD task, we introduce a multimodal graph-conditioned vision-language reconstruction mechanism, which is integrated with graph attention to construct a semantic graph-conditioned reconstruction module (SGCM), this module generates vision-language (VL) tokens through graph-based conditions and enables cross-dimensional interaction between visual and textual features via multihead attention. The reconstructed VL features are then deeply fused using the language vision transformer (LViT), achieving fine-grained feature alignment and high-level semantic interaction. Experimental results on four public datasets demonstrate that MGCR achieves superior performance compared to mainstream CD methods. Our code is available on https://github.com/cn-xvkong/MGCR

遥感变化检测多模态图神经网络视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。