为遥感影像变化描述打造专用指令数据集,提升大模型理解能力
CDChat: A Large Multimodal Model for Remote Sensing Change Description
- 构建遥感变化描述指令数据集,专用于微调多模态模型
- 在双时相遥感图像变化描述任务中表现优于现有方法
- 适合遥感分析、城市监测等需要精准变化识别的场景
大型多模态模型(LMMs)在自然图像领域通过视觉指令微调展现出良好性能,但在遥感图像内容描述任务(如图像或区域定位、分类)中表现不佳。尽管GeoChat尝试改进遥感图像描述,但在双时相遥感图像变化描述这一关键任务上仍存在局限。为此,本文提出一个专门用于变化描述的指令数据集,可用于微调LMM。实验表明,对LLaVA-1.5模型进行少量修改后,在该数据集上微调,能显著提升其在变化描述任务上的表现,验证了该数据集的有效性。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have shown encouraging performance in the natural image domain using visual instruction tuning. However, these LMMs struggle to describe the content of remote sensing images for tasks such as image or region grounding, classification, etc. Recently, GeoChat make an effort to describe the contents of the RS images. Although, GeoChat achieves promising performance for various RS tasks, it struggles to describe the changes between bi-temporal RS images which is a key RS task. This necessitates the development of an LMM that can describe the changes between the bi-temporal RS images. However, there is insufficiency of datasets that can be utilized to tune LMMs. In order to achieve this, we introduce a change description instruction dataset that can be utilized to finetune an LMM and provide better change descriptions for RS images. Furthermore, we show that the LLaVA-1.5 model, with slight modifications, can be finetuned on the change description instruction dataset and achieve favorably better performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。