构建两个遥感图文数据集,提升视觉语言模型零样本泛化能力
A Recipe for Improving Remote Sensing VLM Zero Shot Generalization
- 用谷歌地图地标生成航拍与卫星图的描述,构建新图文数据集
- 在多个公开基准上实现零样本跨模态检索最优性能
- 通过注意力图生成伪标签,增强模型定位能力,适合遥感研究者
基础模型已在诸多AI应用中产生深远影响,推动了以往无法实现的场景。对比视觉语言模型(VLM)在多项任务中表现优异,但在遥感领域应用仍受限,主要因缺乏多样化的遥感视觉-语言数据集。本文提出两个新型图像-标题数据集用于遥感基础模型训练:首个数据集将航拍与卫星图像与由Gemini根据谷歌地图地标生成的描述配对;第二个数据集采用公开网络图片及其对应替代文本,经遥感领域筛选后形成风格与主题更丰富的数据集。这些数据集用于预训练MaMMUT VLM架构,在多个知名公共基准上实现了零样本跨模态检索的最先进性能。此外,我们提出一项持续研究:将VLM对比学习中获得的图像级知识提炼为定位能力。具体通过模型注意力图迭代生成区域伪标签,并用于进一步训练。为缓解噪声注意力图并生成鲁棒分割掩码,引入一种名为Smooth-Attention-Operation的新颖注意力池化机制。
原文摘要 · Abstract (English)
Foundation models have had a significant impact across various AI applications, enabling use cases that were previously impossible. Contrastive Visual Language Models (VLMs), in particular, have outperformed other techniques in many tasks. However, their prevalence in remote sensing (RS) is still limited, due to the scarcity of diverse remote-sensing visual-language datasets. In this work we introduce two novel image-caption datasets for training of remote sensing foundation models. The first dataset pairs aerial and satellite imagery with captions generated by Gemini using landmarks extracted from Google Maps. The second dataset utilizes public web images and their corresponding alt-text, filtered for the remote sensing domain, resulting in a diverse dataset with greater breadth in image styles and subject matter. These datasets are used to pre-train the MaMMUT~\citep{kuo2023mammutsimplearchitecturejoint} VLM architecture, resulting in state-of-the-art generalization performance in zero-shot cross-modal retrieval on well-known public benchmarks. Finally, we present our ongoing research to distill image-level knowledge gained in the VLM contrastive training procedure to enhance the model's localization ability. Specifically, we iteratively generate pseudo-labels for image regions based on the model's attention maps and use these labels for further training. To mitigate noisy attention maps and create robust segmentation masks, we introduce a novel attention-pooling mechanism called the Smooth-Attention-Operation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。