用多视角生成和大模型整合,打造高质量遥感图文数据集
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation
- 分两阶段生成多视角描述,再由大模型融合成完整描述
- 构建21万张遥感图像、130万条标注的高质量数据集
- 新模型仅用4.2%数据就超越现有最优结果,适合遥感研究者
视觉语言基础模型(VLFMs)在遥感(RS)图像中的应用受到广泛关注,因其在多种下游任务中表现出色。然而,高质量、大规模的图像-文本配对训练数据稀缺仍是主要挑战。尽管已有研究构建了大型遥感图文数据集并训练了VLFMs,但因生成标题的方法粗略,数据质量不高,需大量数据才能带来有限性能提升。本文提出一种名为MpGI(多视角生成与整合)的两阶段方法,用于生成高质量遥感图像文本描述。首先,通过规则引导的多模态大语言模型(Rule-MLLM)和多模态大模型(MLLM)生成不同视角的详细描述;随后,利用大语言模型(LLM)将这些多样描述整合为全面连贯的文本。最终,我们构建了包含约21万张遥感图像和130万条文本描述的HQRS-IT-210K数据集。基于该数据集,我们微调了两种VLFMs:CLIP(判别式模型)和CoCa(图像到文本生成模型),分别得到HQRS-CLIP和RS-CoCa模型。实验表明,HQRS-CLIP在多个下游任务中优于此前最先进遥感CLIP模型,且仅使用4.2%的训练数据;RS-CoCa在基准数据集上表现超越其他先进方法,生成的遥感图像描述可媲美甚至超过人工标注。数据集、预训练模型与代码将公开于https://github.com/YiguoHe/HQRS-210K-and-HQRS-CLIP。
原文摘要 · Abstract (English)
The application of Vision-language foundation models (VLFMs) to remote sensing (RS) imagery has garnered significant attention due to their superior capability in various downstream tasks. A key challenge lies in the scarcity of high-quality, large-scale, image-text paired training data. Recently, several works introduced extensive image-text datasets for RS and trained their VLFMs. However, due to the rudimentary methods used for generating captions, the quality of datasets is suboptimal, requiring larger volumes of training data, while only yielding modest performance improvements. In this paper, we propose a two-stage method named MpGI(Multi-Perspective Generation and Integration) for generating high-quality text captions for RS images. Firstly, we generate distinct and detailed descriptions from different perspectives using Rule-MLLM(Multimodal Large Language Model) Relay Generation and MLLMs generation methods. Next, we utilize Large Language Models (LLMs) to integrate these diverse descriptions into comprehensive captions, capturing details from multiple perspectives. Finally, we have created the HQRS-IT-210K dataset, including about 210,000 RS images and 1.3 million captions. We fine-tuned two VLFMs using our dataset: CLIP, a discriminative model, and CoCa, an image-to-text generative model. This process resulted in our proposed HQRS-CLIP and RS-CoCa models. Experimental results demonstrate that HQRS-CLIP surpassed the previous SOTA RS CLIP model in various downstream tasks while using only 4.2\% of the training data. RS-CoCa outperforms other advanced approaches across benchmark datasets and can generate captions for RS images that rival or even exceed manual annotations. Dataset, pre-trained models, and codes will be released at https://github.com/YiguoHe/HQRS-210K-and-HQRS-CLIP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。