arXiv:2409.08582cs.CV2024-09被引 31

首个支持交互式遥感变化分析的多模态模型,可理解用户提问并精准回答变化细节。

ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning

  • 基于多模态指令微调,支持自然语言交互查询
  • 在特定任务上性能媲美或超越现有最优方法
  • 适合需要精准、交互式遥感变化分析的研究者

遥感变化分析对监测地球动态过程至关重要,通过检测时间序列图像中的变化实现。传统变化检测擅长像素级识别,但缺乏上下文理解;近期变化描述生成虽能提供自然语言解释,却无法支持用户定制化交互查询。为此,我们提出ChangeChat,首个专为遥感变化分析设计的双时相视觉-语言模型(VLM)。该模型采用多模态指令微调,可处理复杂查询,包括变化描述、类别量化和变化定位。为提升性能,我们构建了包含87,000条样本的ChangeChat-87k数据集,结合规则生成与GPT辅助标注。实验表明,ChangeChat在多项任务中表现媲美或优于现有最先进方法,并显著超越最新的通用模型GPT-4。代码与预训练权重已开源。

原文摘要 · Abstract (English)

Remote sensing (RS) change analysis is vital for monitoring Earth's dynamic processes by detecting alterations in images over time. Traditional change detection excels at identifying pixel-level changes but lacks the ability to contextualize these alterations. While recent advancements in change captioning offer natural language descriptions of changes, they do not support interactive, user-specific queries. To address these limitations, we introduce ChangeChat, the first bitemporal vision-language model (VLM) designed specifically for RS change analysis. ChangeChat utilizes multimodal instruction tuning, allowing it to handle complex queries such as change captioning, category-specific quantification, and change localization. To enhance the model's performance, we developed the ChangeChat-87k dataset, which was generated using a combination of rule-based methods and GPT-assisted techniques. Experiments show that ChangeChat offers a comprehensive, interactive solution for RS change analysis, achieving performance comparable to or even better than state-of-the-art (SOTA) methods on specific tasks, and significantly surpassing the latest general-domain model, GPT-4. Code and pre-trained weights are available at https://github.com/hanlinwu/ChangeChat.

遥感分析多模态模型交互式分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。