构建大规模遥感变化理解数据集,提升多任务交互式分析能力
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
- 设计双粒度视觉语言模型,融合细粒度空间特征与高层语义
- 在变化描述任务中超越最强基线1.39分(S*m指标)
- 适合遥感智能分析、环境监测方向研究者使用
遥感变化理解(RSCU)对分析遥感图像及人类活动影响环境至关重要。现有数据集在多样化的变化描述、计数和定位任务中缺乏深度理解和交互能力。为此,我们构建了ChangeIMTI——一个包含变化描述、二分类变化检测、变化计数和变化定位四个互补任务的大规模交互式多任务指令数据集。基于此数据集,我们提出一种新型双粒度感知的视觉引导视觉语言模型(ChangeVG),用于双时相遥感图像分析。该模型采用双分支架构,协同融合细粒度空间特征提取与高层语义总结,生成的丰富表征作为辅助提示,指导大视觉语言模型(如Qwen2.5-VL-7B)进行指令微调,促进层次化跨模态学习。我们在四个任务上进行广泛实验,结果表明:在变化描述任务中,本方法在综合评估指标S*m上比最强基线Semantic-CC高出1.39分。此外,通过一系列消融实验验证了方法关键组件的有效性。代码与数据已公开于GitHub。
原文摘要 · Abstract (English)
Remote sensing change understanding (RSCU) is essential for analyzing remote sensing images and understanding how human activities affect the environment. However, existing datasets lack deep understanding and interactions in the diverse change captioning, counting, and localization tasks. To tackle these gaps, we construct ChangeIMTI, a new large-scale interactive multi-task instruction dataset that encompasses four complementary tasks including change captioning, binary change classification, change counting, and change localization. Building upon this new dataset, we further design a novel vision-guided vision-language model (ChangeVG) with dual-granularity awareness for bi-temporal remote sensing images (i.e., two remote sensing images of the same area at different times). The introduced vision-guided module is a dual-branch architecture that synergistically combines fine-grained spatial feature extraction with high-level semantic summarization. These enriched representations further serve as the auxiliary prompts to guide large vision-language models (VLMs) (e.g., Qwen2.5-VL-7B) during instruction tuning, thereby facilitating the hierarchical cross-modal learning. We extensively conduct experiments across four tasks to demonstrate the superiority of our approach. Remarkably, on the change captioning task, our method outperforms the strongest method Semantic-CC by 1.39 points on the comprehensive S*m metric, which integrates the semantic similarity and descriptive accuracy to provide an overall evaluation of change caption. Moreover, we also perform a series of ablation studies to examine the critical components of our method. The source code and associated data for this work are publicly available at Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。