arXiv:2410.17283cs.AI2024-10被引 2

综述遥感视觉语言模型的进展、数据集与提升方法

Advancements in Visual Language Models for Remote Sensing: Datasets, Capabilities, and Enhancement Techniques

  • 梳理遥感领域VLM的核心理论与技术框架
  • 整理多类遥感任务专用数据集及应用场景
  • 分类解析模型优化策略,适合研究者快速入门

近年来,ChatGPT的成功激发了人工智能的新热潮,视觉语言模型(VLM)的发展将这一热情推向新高。与以往将任务建模为判别式模型不同,VLM将任务视为生成式模型,并对齐语言与视觉信息,从而解决更复杂的挑战。遥感(RS)作为高度实用的领域,也采纳了这一趋势,推出了多种基于VLM的遥感方法,展现出良好性能与巨大潜力。本文首先回顾VLM的基础理论,随后总结遥感领域构建的VLM相关数据集及其覆盖的任务类型。最后,根据VLM的核心组件,将改进方法分为三类,并详细介绍了各类方法及其对比。相关项目已发布于https://github.com/taolijie11111/VLMs-in-RS-review。

原文摘要 · Abstract (English)

Recently, the remarkable success of ChatGPT has sparked a renewed wave of interest in artificial intelligence (AI), and the advancements in visual language models (VLMs) have pushed this enthusiasm to new heights. Differring from previous AI approaches that generally formulated different tasks as discriminative models, VLMs frame tasks as generative models and align language with visual information, enabling the handling of more challenging problems. The remote sensing (RS) field, a highly practical domain, has also embraced this new trend and introduced several VLM-based RS methods that have demonstrated promising performance and enormous potential. In this paper, we first review the fundamental theories related to VLM, then summarize the datasets constructed for VLMs in remote sensing and the various tasks they addressed. Finally, we categorize the improvement methods into three main parts according to the core components of VLMs and provide a detailed introduction and comparison of these methods. A project associated with this review has been created at https://github.com/taolijie11111/VLMs-in-RS-review.

视觉语言模型遥感综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。