arXiv:2509.05895cs.CV2025-09中稿 · ICASSP 2026被引 5

用新模型提升遥感影像变化描述准确率

BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model

  • 设计变化提取模块捕捉时序与语义变化
  • 引入提示增强机制,提升空间细节关注
  • 支持双时相变化描述,适合遥感分析场景

双时相卫星影像在城市化监测和灾情评估中具有重要意义。尽管多模态大模型已应用于双时相变化分析,但现有方法通过直接拼接图像对处理,难以充分建模时间相关性和空间语义变化,制约了视觉-语义对齐效果。为此,我们提出BTCChat,一种具备先进双时相变化理解能力的多时相多模态大模型。该模型支持双时相变化描述,同时保留单图理解能力。为更好捕捉图像对中的时序特征与空间语义变化,我们设计了变化提取模块;为进一步增强模型对空间细节的关注,引入提示增强机制,将上下文线索融入提示以提升性能。实验表明,BTCChat在变化描述和视觉问答任务上均达到当前最优表现。代码已公开。

原文摘要 · Abstract (English)

Bi-temporal satellite imagery supports critical applications such as urbanization monitoring and disaster assessment. Although powerful multimodal large language models~(MLLMs) have been applied in bi-temporal change analysis, previous methods process image pairs through direct concatenation, inadequately modeling temporal correlations and spatial semantic changes. This deficiency hampers visual-semantic alignment in change understanding, thereby constraining the overall effectiveness of current approaches. To address this gap, we propose BTCChat, a multi-temporal MLLM with advanced bi-temporal change understanding capability. BTCChat supports bi-temporal change captioning and retains single-image interpretation capability. To better capture temporal features and spatial semantic changes in image pairs, we design a Change Extraction module. Moreover, to enhance the model's attention to spatial details, we introduce a Prompt Augmentation mechanism, which incorporates contextual clues into the prompt to enhance model performance. Experimental results demonstrate that BTCChat achieves state-of-the-art performance on change captioning and visual question answering tasks. The code is available \href{https://github.com/IntelliSensing/BTCChat}{here}.

遥感多模态变化检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。