构建全球多尺度遥感图文数据集,提升视觉语言模型在遥感任务表现。
GAIA: A Global, Multi-modal, Multi-scale Vision-Language Dataset for Remote Sensing Image Analysis
- 从权威来源采集并生成20万+高质量遥感图像-文本对。
- 覆盖近25年全球范围,支持环境变化与灾害分析等应用。
- 适合遥感、AI融合研究者,推动跨模态模型发展。
现有视觉语言模型主要基于网络抓取的噪声图像-文本数据训练,在遥感(RS)领域表现不佳,因常用数据集缺乏科学准确的详细描述,仅关注时间地点等属性。为填补这一空白,我们提出GAIA,一个面向多尺度、多传感器、多模态遥感图像分析的新数据集。GAIA包含201,005对精心筛选的遥感图像-文本对,涵盖多种空间分辨率和遥感模态。不同于以往遥感视觉语言数据集,GAIA聚焦于捕捉环境变迁、自然灾害等动态现象的真实信息,具有全球性与时空均衡分布,覆盖过去25年。数据构建采用两阶段流程:(1)从可信遥感来源定向抓取图像及文本;(2)利用精心设计提示词,借助GPT-4o生成每张图像的五条高质量、科学合理的合成描述。大量实验表明,使用GAIA微调CLIP与BLIP2模型,在遥感图像分类、跨模态检索和图像描述任务上均有显著性能提升。我们已将数据集、自动化处理框架及微调模型权重公开于GitHub:https://github.com/Orion-AI-Lab/GAIA。
原文摘要 · Abstract (English)
Existing Vision-Language Models (VLMs) are predominantly trained on web-scraped, noisy image-text data, exhibiting limited exposure to the specialized domain of RS. This deficiency results in poor performance on RS-specific tasks, as commonly used datasets often lack detailed, scientifically accurate textual descriptions and instead emphasize solely on attributes like date and location. To bridge this critical gap, we introduce GAIA, a novel dataset designed for multi-scale, multi-sensor, and multi-modal RS image analysis. GAIA comprises of 201,005 meticulously curated RS image-text pairs, representing a diverse range of RS modalities associated to different spatial resolutions. Unlike existing vision-language datasets in RS, GAIA specifically focuses on capturing a diverse range of RS applications, providing unique information about environmental changes, natural disasters, and various other dynamic phenomena. The dataset provides a spatially and temporally balanced distribution, spanning across the globe, covering the last 25 years with a balanced temporal distribution of observations. GAIA's construction involved a two-stage process: (1) targeted web-scraping of images and accompanying text from reputable RS-related sources, and (2) generation of five high-quality, scientifically grounded synthetic captions for each image using carefully crafted prompts that leverage the advanced vision-language capabilities of GPT-4o. Our extensive experiments, including fine-tuning of CLIP and BLIP2 models, demonstrate that GAIA significantly improves performance on RS image classification, cross-modal retrieval and image captioning tasks. We make our dataset, automated processing framework and fine-tuned model weights publicly available on our project's GitHub repository: https://github.com/Orion-AI-Lab/GAIA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。