arXiv:2410.06234cs.CVcs.AI2024-10ICLR被引 79

让AI理解地球观测时间序列数据,支持对话式分析变化与场景。

TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

  • 构建可对话的视觉语言模型,处理多时相遥感图像序列。
  • 在变化检测、问答等任务上超越现有模型,零样本表现优异。
  • 适合遥感分析、环境监测等需要时空推理的研究与应用者。

大型视觉语言模型已推动自然图像理解能力的发展。近年来,这些方法被引入地球观测数据领域,但仅限于单张图像输入,难以应对多数真实场景任务。本文提出新型视觉语言助手TEOChat,支持对地球观测时间序列数据进行对话式理解。为训练TEOChat,我们构建了一个指令跟随数据集,涵盖单图与时序任务,包括建筑变化与损毁评估、语义变化检测及时间场景分类。实验表明,TEOChat能完成多样化的空间与时间推理任务,显著优于此前视觉语言模型,甚至达到或超过多个专用任务模型的表现。此外,其在变化检测与问答数据集上展现出色零样本性能,优于GPT-4o和Gemini 1.5 Pro,在多项时序任务中表现更优;在单图任务如场景分类、视觉问答与图像描述方面,也优于同类单图指令模型。相关数据、模型与代码已公开于https://github.com/ermongroup/TEOChat。

原文摘要 · Abstract (English)

Large vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal sequences of earth observation data. To train TEOChat, we curate an instruction-following dataset composed of many single image and temporal tasks including building change and damage assessment, semantic change detection, and temporal scene classification. We show that TEOChat can perform a wide variety of spatial and temporal reasoning tasks, substantially outperforming previous vision and language assistants, and even achieving comparable or better performance than several specialist models trained to perform specific tasks. Furthermore, TEOChat achieves impressive zero-shot performance on a change detection and change question answering dataset, outperforms GPT-4o and Gemini 1.5 Pro on multiple temporal tasks, and exhibits stronger single image capabilities than a comparable single image instruction-following model on scene classification, visual question answering, and captioning. We publicly release our data, model, and code at https://github.com/ermongroup/TEOChat .

遥感分析时序理解多模态对话视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。