无需人工标注,用模型自动生成海量遥感图文数据。
Pushing the Limits of Vision-Language Models in Remote Sensing without Human Annotations
- 用图像解码模型自动构建遥感图文对,省去人工标注
- 生成约960万对数据,显著提升下游任务性能
- 适合遥感、少样本学习等领域的研究者参考
视觉-语言融合基础模型在自然图像领域发展迅速,得益于其丰富且易于网络爬取的图文数据。然而,在遥感领域,尽管已有视觉-语言数据集,其规模仍不足以支撑强健的基础模型。本研究提出一种方法,通过图像解码机器学习模型自动构建视觉-语言数据集,无需人工标注。利用该方法,我们构建了约960万对高分辨率遥感影像的视觉-语言配对数据。所训练模型在零样本分类、语义定位和图文检索等下游任务中表现优于未使用公开视觉-语言数据集的模型。此外,在仅使用视觉编码器的任务如线性探测和k-NN分类中,本模型也优于依赖领域特定视觉-语言数据集的模型。
原文摘要 · Abstract (English)
The prominence of generalized foundation models in vision-language integration has witnessed a surge, given their multifarious applications. Within the natural domain, the procurement of vision-language datasets to construct these foundation models is facilitated by their abundant availability and the ease of web crawling. Conversely, in the remote sensing domain, although vision-language datasets exist, their volume is suboptimal for constructing robust foundation models. This study introduces an approach to curate vision-language datasets by employing an image decoding machine learning model, negating the need for human-annotated labels. Utilizing this methodology, we amassed approximately 9.6 million vision-language paired datasets in VHR imagery. The resultant model outperformed counterparts that did not leverage publicly available vision-language datasets, particularly in downstream tasks such as zero-shot classification, semantic localization, and image-text retrieval. Moreover, in tasks exclusively employing vision encoders, such as linear probing and k-NN classification, our model demonstrated superior efficacy compared to those relying on domain-specific vision-language datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。