首个面向遥感变化描述的多模态大模型,提升细粒度变化理解能力。
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning

- 构建数据生成范式,释放指令数据集RSICI
- 7B参数模型超越更大规模模型表现
- 适合遥感变化分析、地理监测等应用
遥感图像变化描述(RSICC)旨在刻画双时相遥感图像间的差异,具有重要研究与应用价值。现有方法多依赖传统深度学习架构,模型容量有限制约性能。尽管大模型后训练技术在通用领域取得成功,但其直接迁移至RSICC面临数据稀缺与细粒度变化理解需求的挑战。为此,我们提出首个面向RSICC的大视觉-语言模型后训练框架RSICCLLM。具体而言,设计数据生成范式,发布指令数据集RSICI,建立任务专用基准。引入差异感知监督微调,显式提取变化表征并引导模型感知时间差异。此外,提出双负样本偏好优化(DNPO),采用两种互补的负样本构造策略构建偏好数据集RSICP,进一步优化模型性能。大量实验验证了RSICCLLM的卓越能力:仅用7B参数即超越显著更大规模模型。代码与数据集将公开于https://github.com/keaill/RSICCLLM。
原文摘要 · Abstract (English)
Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, most existing methods rely on conventional deep learning architectures, and the limited model capacity constrains performance. Although large-model post-training techniques have achieved great success in general domains, their direct transfer to RSICC remains challenging due to data scarcity and the need for fine-grained change understanding. To address this, we propose RSICCLLM, the first post-training framework for large vision-language models in RSICC. Specifically, we design a data generation paradigm, release the instruction dataset RSICI, and establish a task-specific RSICC benchmark. We further introduce Difference-aware Supervised Fine-tuning to explicitly extract change representations and guide the model in perceiving and understanding temporal differences. In addition, we propose Dual-Negative Preference Optimization (DNPO), which employs two complementary negative-sample construction strategies to construct the preference dataset RSICP and further refine model performance. Extensive experiments validate the superior capability of RSICCLLM, which achieves outstanding results with only 7B parameters, surpassing models of substantially larger scales. The code and dataset will be made publicly available at https://github.com/keaill/RSICCLLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。