arXiv:2509.24875cs.CVcs.LG2025-09

用环境信息控制生成卫星图,更准更稳。

Environment-Aware Satellite Image Generation with Diffusion Models

  • 通过文本、元数据和视觉数据三重控制生成卫星图像
  • 在6项指标上优于现有方法,对缺失数据更鲁棒
  • 首个融合三模态的公开卫星数据集,适合遥感应用

基于扩散模型的生成模型因能生成高质量图像而受到关注。尽管应用到遥感领域尚不直接,但最近尝试已初步利用大量包含多模态信息的公开数据集。现有方法存在局限:依赖有限环境上下文,难以处理缺失或损坏数据,且生成结果常无法准确反映用户意图。本文提出一种新型扩散模型,可同时基于三种控制信号生成卫星图像:文本、元数据和视觉数据。相比以往工作,本方法首次将动态环境条件作为控制信号,并采用元数据融合策略建模属性嵌入间的交互关系,以应对部分缺失或损坏的数据。实验表明,该方法在单图生成与时序生成任务中均表现更优:定性上对缺失元数据更鲁棒,对输入响应更灵敏;定量上在6项指标上均取得更高保真度、准确率与质量。结果支持了环境上下文有助于提升卫星图像生成性能的假设,使本模型成为下游任务的有力候选。所构建的三模态数据集为首个公开可用的此类数据集。

原文摘要 · Abstract (English)

Diffusion-based foundation models have recently garnered much attention in the field of generative modeling due to their ability to generate images of high quality and fidelity. Although not straightforward, their recent application to the field of remote sensing signaled the first successful trials towards harnessing the large volume of publicly available datasets containing multimodal information. Despite their success, existing methods face considerable limitations: they rely on limited environmental context, struggle with missing or corrupted data, and often fail to reliably reflect user intentions in generated outputs. In this work, we propose a novel diffusion model conditioned on environmental context, that is able to generate satellite images by conditioning from any combination of three different control signals: a) text, b) metadata, and c) visual data. In contrast to previous works, the proposed method is i) to our knowledge, the first of its kind to condition satellite image generation on dynamic environmental conditions as part of its control signals, and ii) incorporating a metadata fusion strategy that models attribute embedding interactions to account for partially corrupt and/or missing observations. Our method outperforms previous methods both qualitatively (robustness to missing metadata, higher responsiveness to control inputs) and quantitatively (higher fidelity, accuracy, and quality of generations measured using 6 different metrics) in the trials of single-image and temporal generation. The reported results support our hypothesis that conditioning on environmental context can improve the performance of foundation models for satellite imagery, and render our model a promising candidate for usage in downstream tasks. The collected 3-modal dataset is to our knowledge, the first publicly-available dataset to combine data from these three different mediums.

卫星图像扩散模型多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。