打造全球遥感图像生成新范式,支持任意分辨率与场景的文本驱动生成。
Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model
- 基于扩散模型构建13亿参数基础模型,融合分辨率引导机制。
- 在1050万对遥感图文数据上训练,生成效果超越此前模型26.23点FID。
- 适合遥感分析、环境监测等需要灵活生成大范围图像的科研人员。
生成式基础模型已在自然图像领域实现大规模文本驱动生成,但在遥感领域仍缺乏相关研究。现有遥感图像-文本数据集规模小、地理覆盖有限且场景类型单一,现有文本到图像生成方法难以实现全局尺度、多分辨率可控、无边界图像生成。为此,本文提出两大贡献:全球尺度的Git-10M数据集与Text2Earth基础模型。Git-10M包含1050万张图像-文本对,规模为此前最大数据集的5倍,涵盖广泛地理场景并附带分辨率信息,显著提升数据量与多样性。在此基础上,我们提出基于扩散框架的13亿参数生成模型Text2Earth,集成分辨率引导机制,支持用户指定图像分辨率;采用动态条件适配策略,提升训练与推理阶段图像质量。Text2Earth在零样本文本到图像生成中表现优异,具备强大泛化能力,适用于无边界场景构建、图像编辑与跨模态生成等任务。其性能超越以往受限于固定尺寸与有限场景的模型,在旧基准数据集上达到+26.23 FID和+20.95%零样本分类准确率提升。
原文摘要 · Abstract (English)
Generative foundation models have advanced large-scale text-driven natural image generation, becoming a prominent research trend across various vertical domains. However, in the remote sensing field, there is still a lack of research on large-scale text-to-image (text2image) generation technology. Existing remote sensing image-text datasets are small in scale and confined to specific geographic areas and scene types. Besides, existing text2image methods have struggled to achieve global-scale, multi-resolution controllable, and unbounded image generation. To address these challenges, this paper presents two key contributions: the Git-10M dataset and the Text2Earth foundation model. Git-10M is a global-scale image-text dataset comprising 10.5 million image-text pairs, 5 times larger than the previous largest one. The dataset covers a wide range of geographic scenes and contains resolution information, significantly surpassing existing datasets in both size and diversity. Building on Git-10M, we propose Text2Earth, a 1.3 billion parameter generative foundation model based on the diffusion framework to model global-scale remote sensing scenes. Text2Earth integrates a resolution guidance mechanism, enabling users to specify image resolutions. A dynamic condition adaptation strategy is proposed for training and inference to improve image quality. Text2Earth excels in zero-shot text2image generation and demonstrates robust generalization and flexibility across multiple tasks, including unbounded scene construction, image editing, and cross-modal image generation. This robust capability surpasses previous models restricted to the basic fixed size and limited scene types. On the previous benchmark dataset, Text2Earth outperforms previous models with an improvement of +26.23 FID and +20.95% Zero-shot Cls-OA metric.Our project page is https://chen-yang-liu.github.io/Text2Earth
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。