首个基于多光谱遥感数据的视觉语言模型,提升地球观测理解能力。
Beyond the Visible: Multispectral Vision-Language Learning for Earth Observation
- 用多光谱数据预训练视觉语言模型,突破传统仅依赖可见光的局限。
- 在零样本分类和检索任务中,准确率提升6.77%,mAP提高4.63%。
- 适合遥感、环境监测领域研究者,推动多源数据融合应用。
面向地球观测(EO)的视觉语言模型通常仅使用可见光图像作为输入,未能充分利用卫星记录的多光谱通道中的丰富信息。为此,我们提出Llama3-MS-CLIP,首个在大规模多光谱数据集上通过对比学习预训练的视觉语言模型,并报告了扩展光谱范围带来的性能提升。同时,我们构建了迄今为止最大的多光谱图像-文本数据集,包含一百万条哨兵-2(Sentinel-2)样本及其由Llama3-LLaVA-Next和Overture Maps生成的文本描述。我们开发了一套可扩展的图文生成流水线,并经领域专家验证。在三个不同复杂度的数据集上评估了Llama3-MS-CLIP的多光谱零样本图像分类与检索性能。结果表明,该模型显著优于其他基于RGB的方法,在分类准确率上平均提升6.77%,检索mAP提升4.63%。这些结果凸显了多光谱视觉语言学习的重要性。相关数据集、代码及模型权重已公开于https://github.com/IBM/MS-CLIP。
原文摘要 · Abstract (English)
Vision-language models for Earth observation (EO) typically rely on the visual spectrum of data as the only model input, thus failing to leverage the rich spectral information available in the multispectral channels recorded by satellites. Therefore, we introduce Llama3-MS-CLIP, the first vision-language model pre-trained with contrastive learning on a large-scale multispectral dataset and report on the performance gains due to the extended spectral range. Furthermore, we present the largest-to-date image-caption dataset for multispectral data, consisting of one million Sentinel-2 samples and corresponding textual descriptions generated using Llama3-LLaVA-Next and Overture Maps data. We develop a scalable captioning pipeline, which is validated by domain experts. We evaluate Llama3-MS-CLIP on multispectral zero-shot image classification and retrieval using three datasets of varying complexity. Our results demonstrate that Llama3-MS-CLIP significantly outperforms other RGB-based approaches, improving classification accuracy by +6.77% on average and retrieval performance by +4.63% mAP compared to the second-best model. Our results emphasize the relevance of multispectral vision-language learning. The image-caption dataset, code, and model weights are available at https://github.com/IBM/MS-CLIP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。