用多模态大模型实现遥感变化理解,突破时间感知与定位精度瓶颈。
Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models

- 设计新架构Delta-LLaVA,通过变化增强注意力与因果注意力解决时序混淆问题。
- 在18万样本的Delta-QA基准上,变化推理准确率显著超越现有模型。
- 适合遥感智能分析、环境监测等需要精准时空变化理解的场景。
尽管多模态大语言模型在通用视觉语言任务中表现优异,但在遥感变化理解方面受限于固有的“时间盲区”。现有架构缺乏多时相对比推理机制,且难以实现精确空间定位。为此,本文首次提出Delta-QA,一个包含18万样本的视觉问答基准,涵盖双时相与三时相场景,将变化解释划分为四个渐进认知维度。方法上,提出针对多时相遥感理解的Delta-LLaVA框架,通过三项核心创新克服简单特征拼接的局限:变化增强注意力模块系统性分离并放大视觉差异;变化分割模块利用变化先验嵌入提取可微差分特征输入大模型;局部因果注意力防止跨时相上下文泄露。大量实验表明,Delta-LLaVA在复杂变化推断与高精度边界定位任务中显著优于主流通用大模型与专用分割模型,建立了统一的地球观测智能分析框架。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hindered by a fundamental "temporal blindness". Existing architectures lack intrinsic mechanisms for multi-temporal contrastive reasoning and struggle with precise spatial grounding. To address this, we first introduce Delta-QA, a comprehensive benchmark comprising 180k visual question-answering samples. Delta-QA unifies pixel-level segmentation and visual question answering across bi- and tri-temporal scenarios, structuring change interpretation into four progressive cognitive dimensions. Methodologically, we propose Delta-LLaVA, a novel MLLM framework explicitly tailored for multi-temporal remote sensing interpretation. It overcomes the limitations of naive feature concatenation through three core innovations: a Change-Enhanced Attention module that systematically isolates and amplifies visual differences, a Change-SEG module utilizing Change Prior Embedding to extract differentiable difference features as input for the LLM, and Local Causal Attention to prevent cross-temporal contextual leakage. Extensive experiments demonstrate that Delta-LLaVA decisively outperforms leading generalist MLLMs and specialized segmentation models in complex change deduction and high-precision boundary localization, establishing a unified framework for earth observation intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。