构建首个多模态多尺度遥感图文生成数据集,解决遥感图像生成难题。
MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation
- 整合9个公开遥感数据集,标准化处理并生成210万组图文对。
- 通过预训练视觉语言模型生成文本提示,支持多场景、多分辨率图像生成。
- 适合遥感图像生成、扩散模型应用与跨模态研究者使用。
近年来,基于扩散的生成范式凭借精准的分布建模和稳定的训练过程,在文本驱动的通用图像生成方面取得了显著成果。然而,由于缺乏涵盖多种模态、地面采样距离(GSD)和场景的综合性遥感图像生成数据集,生成在尺度和视角上与通用图像差异巨大的遥感图像仍面临巨大挑战。本文提出一个面向多样化遥感场景的多模态、多GSD、多场景遥感(MMM-RS)数据集与基准测试。我们首先收集了9个公开的遥感数据集,并对所有样本进行标准化处理。为实现遥感图像与文本语义信息的关联,利用大规模预训练视觉-语言模型自动生成文本提示,并辅以人工修正,形成信息丰富的图文对(包括多模态图像)。特别地,设计方法在同一样本中生成不同GSD及各种环境(如低光、雾天)的图像。经过大量人工筛选与标注修正,最终构建包含约210万组文本-图像对的MMM-RS数据集。大量实验验证,该数据集可使现成的扩散模型在多种模态、场景、天气条件和GSD下生成多样化的遥感图像。数据集已开源:https://github.com/ljl5261/MMM-RS。
原文摘要 · Abstract (English)
Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general images in terms of scale and perspective remains a formidable challenge due to the lack of a comprehensive remote sensing image generation dataset with various modalities, ground sample distances (GSD), and scenes. In this paper, we propose a Multi-modal, Multi-GSD, Multi-scene Remote Sensing (MMM-RS) dataset and benchmark for text-to-image generation in diverse remote sensing scenarios. Specifically, we first collect nine publicly available RS datasets and conduct standardization for all samples. To bridge RS images to textual semantic information, we utilize a large-scale pretrained vision-language model to automatically output text prompts and perform hand-crafted rectification, resulting in information-rich text-image pairs (including multi-modal images). In particular, we design some methods to obtain the images with different GSD and various environments (e.g., low-light, foggy) in a single sample. With extensive manual screening and refining annotations, we ultimately obtain a MMM-RS dataset that comprises approximately 2.1 million text-image pairs. Extensive experimental results verify that our proposed MMM-RS dataset allows off-the-shelf diffusion models to generate diverse RS images across various modalities, scenes, weather conditions, and GSD. The dataset is available at https://github.com/ljl5261/MMM-RS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。