arXiv:2412.02573cs.CV2024-12综述被引 49

遥感时空视觉语言模型让机器读懂地球变化,还能用自然语言描述和问答。

Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey

  • 融合视觉与语言模态,实现时空变化的语义理解
  • 支持变化描述生成、问答与定位,突破传统二值掩码限制
  • 首次系统综述该领域模型演进与数据评估体系

多时相遥感图像的解读对监测地球动态过程至关重要,但以往的变化检测方法仅输出二值或语义掩码,难以提供人类可读的洞察。近年来,视觉语言模型(VLMs)通过融合视觉与语言模态,开启了新方向:实现时空视觉语言理解——不仅能捕捉空间与时间依赖关系以识别变化,还能对时序图像进行更丰富的交互式语义分析(如生成描述性标题、回答自然语言问题)。本文首次全面综述遥感时空视觉语言模型(RS-STVLMs)的发展。涵盖从早期任务特定模型到近年利用强大大语言模型的通用基础模型的演进历程。讨论代表性任务进展,包括变化描述生成、变化问答与变化定位。系统剖析模型核心组件与关键技术,并回顾推动该领域的数据集与评估指标。通过任务层面洞察与架构模式深度分析,旨在揭示当前成就并规划未来研究方向。我们将持续追踪相关工作:https://github.com/Chen-Yang-Liu/Awesome-RS-SpatioTemporal-VLMs

原文摘要 · Abstract (English)

The interpretation of multi-temporal remote sensing imagery is critical for monitoring Earth's dynamic processes-yet previous change detection methods, which produce binary or semantic masks, fall short of providing human-readable insights into changes. Recent advances in Vision-Language Models (VLMs) have opened a new frontier by fusing visual and linguistic modalities, enabling spatio-temporal vision-language understanding: models that not only capture spatial and temporal dependencies to recognize changes but also provide a richer interactive semantic analysis of temporal images (e.g., generate descriptive captions and answer natural-language queries). In this survey, we present the first comprehensive review of RS-STVLMs. The survey covers the evolution of models from early task-specific models to recent general foundation models that leverage powerful large language models. We discuss progress in representative tasks, such as change captioning, change question answering, and change grounding. Moreover, we systematically dissect the fundamental components and key technologies underlying these models, and review the datasets and evaluation metrics that have driven the field. By synthesizing task-level insights with a deep dive into shared architectural patterns, we aim to illuminate current achievements and chart promising directions for future research in spatio-temporal vision-language understanding for remote sensing. We will keep tracing related works at https://github.com/Chen-Yang-Liu/Awesome-RS-SpatioTemporal-VLMs

遥感视觉语言模型时空理解综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。