融合对比学习与掩码建模,提升遥感图像的视觉语言对齐能力。
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
- 结合对比学习与掩码建模,实现多模态对齐与视觉表征优化。
- 在SpaceNet1数据集上,语义分割mIOU提升6%,优于SkyCLIP基准。
- 兼顾零样本分类能力,适合需要跨模态理解的遥感应用。
遥感图像富含物体与上下文视觉信息。近期趋势是利用配对的卫星图像与文本描述进行预训练,以构建下游任务表现优异的编码器。然而,尽管对比图像-文本方法(如CLIP)能实现视觉-语言对齐并支持零样本分类,其纯视觉任务性能通常低于仅图像预训练方法(如MAE)。本文提出FLAVARS,一种结合对比学习与掩码建模,并通过对比位置编码实现地理空间对齐的预训练方法。实验表明,FLAVARS在视觉任务(如KNN分类和语义分割)上显著优于基线方法SkyCLIP,于SpaceNet1数据集上提升6% mIOU,同时保留了零样本分类能力,而传统MAE预训练方法则不具备此特性。
原文摘要 · Abstract (English)
Remote sensing imagery is dense with objects and contextual visual information. There is a recent trend to combine paired satellite images and text captions for pretraining performant encoders for downstream tasks. However, while contrastive image-text methods like CLIP enable vision-language alignment and zero-shot classification ability, vision-only downstream performance tends to degrade compared to image-only pretraining, such as MAE. In this paper, we propose FLAVARS, a pretraining method that combines the best of both contrastive learning and masked modeling, along with geospatial alignment via contrastive location encoding. We find that FLAVARS significantly outperforms a baseline of SkyCLIP for vision-only tasks such as KNN classification and semantic segmentation, +6\% mIOU on SpaceNet1, while retaining the ability to perform zero-shot classification, unlike MAE pretrained methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。