arXiv:2504.08548cs.GRcs.CV2025-04CVPR被引 4

用扩散模型统一生成多源遥感影像,支持任意模态间零样本转换。

COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails

  • 基于序列扩散变换器,每种模态独立控制时间步嵌入。
  • 在Major TOM数据集上生成高质量遥感缩略图,跨模态生成效果佳。
  • 适合需要多源遥感数据生成与迁移的研究者使用。

在遥感领域,来自不同传感器的多模态数据可提供丰富信息,但跨模态学习统一表示仍是重大挑战。传统方法多局限于单模态或双模态处理。本文提出COP-GEN-Beta,一个在Major TOM数据集的光学、雷达和高程数据上训练的生成式扩散模型。该模型能够将任意模态子集映射到任意其他模态,实现训练后零样本模态转换。其核心为序列式扩散变换器,每种模态由独立的时间步嵌入控制。我们在Major TOM数据集的缩略图上进行了全面评估,验证了模型生成高质量样本的能力。定性与定量分析均表明其性能优异,具备作为未来遥感任务强大预训练模型的潜力。

原文摘要 · Abstract (English)

In remote sensing, multi-modal data from various sensors capturing the same scene offers rich opportunities, but learning a unified representation across these modalities remains a significant challenge. Traditional methods have often been limited to single or dual-modality approaches. In this paper, we introduce COP-GEN-Beta, a generative diffusion model trained on optical, radar, and elevation data from the Major TOM dataset. What sets COP-GEN-Beta apart is its ability to map any subset of modalities to any other, enabling zero-shot modality translation after training. This is achieved through a sequence-based diffusion transformer, where each modality is controlled by its own timestep embedding. We extensively evaluate COP-GEN-Beta on thumbnail images from the Major TOM dataset, demonstrating its effectiveness in generating high-quality samples. Qualitative and quantitative evaluations validate the model's performance, highlighting its potential as a powerful pre-trained model for future remote sensing tasks.

遥感生成模型扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。