统一生成五模态遥感图像,实现任意模态间互转。
MetaEarth-MM: Unified Multimodal Remote Sensing Image Generation with Scene-centered Joint Modeling

- 以场景为中心建模,先提取隐含场景表示再生成目标模态。
- 在280万张全球多分辨率图像上训练,支持任意模态对转换。
- 适合遥感多模态数据生成与跨模态分析的研究者使用。
多模态遥感图像对地球观测至关重要,但实际中完整配对数据常稀缺。现有生成方法多采用孤立的成对模态转换,随模态和任务增加,泛化性与可扩展性受限。本文提出生成式基础模型MetaEarth-MM,可在统一框架内实现五种模态间的联合生成与任意模态互转。基于多模态观测内在场景一致性,引入场景中心联合建模范式:不依赖直接外观映射,而是先从已有观测中推断隐含场景表示,再据此生成目标模态。为支持训练,构建了包含280万张多分辨率全球图像及220万对齐样本的EarthMM大规模数据集。大量实验表明,MetaEarth-MM不仅具备强生成能力与跨任务鲁棒泛化性,还在数据与表征层面支持下游任务,展现出作为跨模态地球观测通用基础模型的潜力。代码与数据集将开源。
原文摘要 · Abstract (English)
Multi-modal remote sensing images are vital for Earth observation, yet complete paired observations are often scarce in practice. Existing generative methods commonly address this problem through isolated pairwise modality translation, but their versatility and scalability remain limited as the number of modalities and generation tasks increases. Here, we develop a generative foundation model MetaEarth-MM for multi-modal remote sensing imagery, enabling paired joint generation and any-to-any translation across five modalities within a unified model. Recognizing the intrinsic scene consistency underlying multi-modal observations, we introduce a scene-centered joint modeling paradigm in MetaEarth-MM. Unlike previous methods that rely on direct appearance-level cross-modal mapping, our model organizes the generation around the underlying scene content. Specifically, MetaEarth-MM adopts a decoupled architecture that first infers a latent scene representation from available observations, and then generates target modalities conditioned on this intermediate state. To support training, we further construct EarthMM, a large-scale dataset comprising 2.8 million multi-resolution global images with 2.2 million aligned pairs. Extensive experiments demonstrate that MetaEarth-MM not only exhibits strong generative capability and robust generalization across diverse generation tasks, but also supports downstream tasks at both data and representation levels, highlighting its potential as a general foundation model for cross-modal Earth observation. The code and dataset will be available at https://github.com/YZPioneer/MetaEarth-MM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。