arXiv:2410.23370cs.CV2024-10中稿 · ACM SIGSPATIAL 202…被引 15

多语言遥感视觉语言模型,提升跨模态检索与零样本分类性能。

Multilingual Vision-Language Pre-training for the Remote Sensing Domain

  • 基于多语言CLIP,融合自监督局部全局对齐与标准对比学习。
  • 在9种语言的翻译数据上训练,英文任务性能也显著提升。
  • 适用于多语言遥感图像文本检索与零样本分类场景。

当前基于对比语言-图像预训练(CLIP)的方法广泛用于遥感领域的视觉语言任务,如跨模态检索。现有方法主要依赖在人类标注的图像-标题数据集上微调模型,或使用其他遥感图像标注生成的合成图像-标题对。然而,预训练机制的探索较少,且极少考虑多语言输入。本文提出一种新型遥感领域视觉-语言模型,通过微调多语言CLIP,并结合基于局部与全局表示对齐的自监督方法,辅以标准对比学习目标。模型训练基于已有的遥感图像-英文标题数据集,随后通过自动机器翻译生成九种额外语言的标题。实验表明,翻译数据有效提升了性能,甚至改善了英文任务表现。所提出的模型命名为遥感多语言CLIP(RS-M-CLIP),在多种视觉语言任务中达到最优结果,包括跨模态和多语言图像-文本检索、零样本图像分类。

原文摘要 · Abstract (English)

Methods based on Contrastive Language-Image Pre-training (CLIP) are nowadays extensively used in support of vision-and-language tasks involving remote sensing data, such as cross-modal retrieval. The adaptation of CLIP to this specific domain has relied on model fine-tuning with the standard contrastive objective, using existing human-labeled image-caption datasets, or using synthetic data corresponding to image-caption pairs derived from other annotations over remote sensing images (e.g., object classes). The use of different pre-training mechanisms has received less attention, and only a few exceptions have considered multilingual inputs. This work proposes a novel vision-and-language model for the remote sensing domain, exploring the fine-tuning of a multilingual CLIP model and testing the use of a self-supervised method based on aligning local and global representations from individual input images, together with the standard CLIP objective. Model training relied on assembling pre-existing datasets of remote sensing images paired with English captions, followed by the use of automated machine translation into nine additional languages. We show that translated data is indeed helpful, e.g. improving performance also on English. Our resulting model, which we named Remote Sensing Multilingual CLIP (RS-M-CLIP), obtains state-of-the-art results in a variety of vision-and-language tasks, including cross-modal and multilingual image-text retrieval, or zero-shot image classification.

遥感多语言CLIP跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。