提出新模型与数据集,显著提升遥感图像变化描述生成能力
CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset
- 设计差异感知融合模块,增强多时相图像差分特征表达
- 构建含20万图像对的CC-Foundation数据集,支持持续预训练
- 三阶段渐进训练确保模型深度融合,性能超越现有方法
遥感图像变化描述(RSICC)旨在生成多时相遥感图像中地表变化的自然语言描述,涵盖变化对象的类别、位置及动态(如新增或消失)。现有方法虽尝试利用多模态大模型(MLLM)的长序列理解与推理能力,但因缺乏充分数据支持,常破坏MLLM的原始特征传输路径,干扰其内在知识,限制在该任务中的潜力。本文提出新模型CCExpert,基于先进多模态大模型框架:首先设计差异感知融合模块,捕捉双时相图像的多尺度差异并融入原图上下文,提升差分特征信噪比;其次构建高质量、多样化的数据集CC-Foundation,包含20万图像对和120万条描述,为领域持续预训练提供数据支撑;最后采用三阶段渐进训练流程,确保差异感知模块与预训练MLLM深度集成。CCExpert在LEVIR-CC基准上取得$S^*_m=81.80$的显著性能,大幅超越此前最优方法。代码与部分数据集将开源。
原文摘要 · Abstract (English)
Remote Sensing Image Change Captioning (RSICC) aims to generate natural language descriptions of surface changes between multi-temporal remote sensing images, detailing the categories, locations, and dynamics of changed objects (e.g., additions or disappearances). Many current methods attempt to leverage the long-sequence understanding and reasoning capabilities of multimodal large language models (MLLMs) for this task. However, without comprehensive data support, these approaches often alter the essential feature transmission pathways of MLLMs, disrupting the intrinsic knowledge within the models and limiting their potential in RSICC. In this paper, we propose a novel model, CCExpert, based on a new, advanced multimodal large model framework. Firstly, we design a difference-aware integration module to capture multi-scale differences between bi-temporal images and incorporate them into the original image context, thereby enhancing the signal-to-noise ratio of differential features. Secondly, we constructed a high-quality, diversified dataset called CC-Foundation, containing 200,000 image pairs and 1.2 million captions, to provide substantial data support for continue pretraining in this domain. Lastly, we employed a three-stage progressive training process to ensure the deep integration of the difference-aware integration module with the pretrained MLLM. CCExpert achieved a notable performance of $S^*_m=81.80$ on the LEVIR-CC benchmark, significantly surpassing previous state-of-the-art methods. The code and part of the dataset will soon be open-sourced at https://github.com/Meize0729/CCExpert.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。