综述遥感视觉语言模型的进展与未来方向。
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- 按对比学习、指令微调、文本生成三类梳理遥感VLM方法
- 系统总结主流模型架构与训练策略,涵盖跨模态对齐能力
- 适合关注遥感智能分析与多模态融合的研究者阅读
视觉语言建模(VLM)旨在弥合图像与自然语言之间的信息鸿沟。在先在海量图文对上预训练、再在特定任务数据上微调的新范式下,遥感领域的VLM已取得显著进展。这些模型得益于广泛通用知识的吸收,在多种遥感数据分析任务中表现出色,并具备与用户进行对话交互的能力。本文旨在为遥感领域提供一份及时且全面的VLM发展综述。首先,我们构建了遥感VLM的分类体系:对比学习、视觉指令微调和文本条件图像生成。针对每类方法,详述常用网络架构与预训练目标。其次,全面回顾现有工作,涵盖基于对比的VLM中的基础模型与任务适配方法,指令型VLM中的架构升级、训练策略与模型能力,以及生成式基础模型及其代表性下游应用。第三,总结用于VLM预训练、微调与评估的数据集,分析其构建方法(包括图像来源与标题生成方式)及关键属性,如规模与任务适应性。最后,展望未来研究方向:跨模态表征对齐、模糊需求理解、解释驱动的模型可靠性、持续可扩展的模型能力,以及包含更丰富模态与更大挑战的大规模数据集。
原文摘要 · Abstract (English)
Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote sensing domain has made significant progress. The resulting models benefit from the absorption of extensive general knowledge and demonstrate strong performance across a variety of remote sensing data analysis tasks. Moreover, they are capable of interacting with users in a conversational manner. In this paper, we aim to provide the remote sensing community with a timely and comprehensive review of the developments in VLM using the two-stage paradigm. Specifically, we first cover a taxonomy of VLM in remote sensing: contrastive learning, visual instruction tuning, and text-conditioned image generation. For each category, we detail the commonly used network architecture and pre-training objectives. Second, we conduct a thorough review of existing works, examining foundation models and task-specific adaptation methods in contrastive-based VLM, architectural upgrades, training strategies and model capabilities in instruction-based VLM, as well as generative foundation models with their representative downstream applications. Third, we summarize datasets used for VLM pre-training, fine-tuning, and evaluation, with an analysis of their construction methodologies (including image sources and caption generation) and key properties, such as scale and task adaptability. Finally, we conclude this survey with insights and discussions on future research directions: cross-modal representation alignment, vague requirement comprehension, explanation-driven model reliability, continually scalable model capabilities, and large-scale datasets featuring richer modalities and greater challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。