arXiv:2507.22346cs.CV2025-07被引 7

让卫星图像变化分析可交互问答,像聊天一样查变化。

DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception

  • 用指令引导的视觉差分感知模块,理解多时相影像差异。
  • 在10.5万条指令数据上训练,支持多轮问答和定量分析。
  • 适合遥感、城市规划等需要动态变化追踪的研究者。

准确解读多时相卫星影像中的地表覆盖变化对现实场景至关重要。现有方法通常仅提供单次变化掩码或静态描述,难以支持交互式、按需查询的分析。本文提出遥感图像变化分析(RSICA)新范式,融合变化检测与视觉问答能力,实现对双时相遥感图像的多轮、指令引导式变化探索。为此,我们构建了包含10.5万条样本的大型指令跟随数据集ChangeChat-105k,涵盖六类交互:变化描述、分类、量化、定位、开放式问答及多轮对话。基于此,提出DeltaVLM端到端架构,具备三项创新:(1) 经微调的双时相视觉编码器以捕捉时间差异;(2) 带跨语义关系度量(CSRM)机制的视觉差分感知模块以解析变化;(3) 指令引导的Q-former模块,从视觉变化中精准提取与文本指令相关的差异信息。模型在ChangeChat-105k上训练,仅优化视觉与对齐模块,保持大语言模型冻结,提升效率。大量实验与消融研究证明,DeltaVLM在单轮描述与多轮交互分析上均达到当前最优性能,优于现有多模态大模型与遥感视觉语言模型。代码、数据集及预训练权重已公开于https://github.com/hanlinwu/DeltaVLM。

原文摘要 · Abstract (English)

Accurate interpretation of land-cover changes in multi-temporal satellite imagery is critical for real-world scenarios. However, existing methods typically provide only one-shot change masks or static captions, limiting their ability to support interactive, query-driven analysis. In this work, we introduce remote sensing image change analysis (RSICA) as a new paradigm that combines the strengths of change detection and visual question answering to enable multi-turn, instruction-guided exploration of changes in bi-temporal remote sensing images. To support this task, we construct ChangeChat-105k, a large-scale instruction-following dataset, generated through a hybrid rule-based and GPT-assisted process, covering six interaction types: change captioning, classification, quantification, localization, open-ended question answering, and multi-turn dialogues. Building on this dataset, we propose DeltaVLM, an end-to-end architecture tailored for interactive RSICA. DeltaVLM features three innovations: (1) a fine-tuned bi-temporal vision encoder to capture temporal differences; (2) a visual difference perception module with a cross-semantic relation measuring (CSRM) mechanism to interpret changes; and (3) an instruction-guided Q-former to effectively extract query-relevant difference information from visual changes, aligning them with textual instructions. We train DeltaVLM on ChangeChat-105k using a frozen large language model, adapting only the vision and alignment modules to optimize efficiency. Extensive experiments and ablation studies demonstrate that DeltaVLM achieves state-of-the-art performance on both single-turn captioning and multi-turn interactive change analysis, outperforming existing multimodal large language models and remote sensing vision-language models. Code, dataset and pre-trained weights are available at https://github.com/hanlinwu/DeltaVLM.

遥感分析交互式生成视觉问答变化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。