用大规模模拟数据预训练,让遥感图像融合模型跨传感器通用。
Leveraging Large-Scale Pretrained Spatial-Spectral Priors for General Zero-Shot Pansharpening
- 用自然图像和遥感图加多种退化操作生成模拟数据,预训练空间光谱先验。
- 零样本下在6个卫星数据集上表现优异,一样本微调仅需少量真实数据。
- 适用于不同网络结构,为跨域遥感融合提供新范式,适合遥感研究者。
现有深度学习遥感图像融合方法因真实训练数据有限及不同卫星传感器间域差异,泛化能力差。本文提出一种新预训练策略,利用大规模模拟数据学习鲁棒的空间-光谱先验。具体地,通过在ImageNet和SkyScript遥感图像上施加模糊、噪声、下采样等退化操作,以及波段生成、通道打乱、高通滤波、色彩抖动等多种增强,构建多样化模拟数据集,并在此基础上预训练融合模型,以学习可迁移的空间-光谱表示。预训练模型在六个数据集(WorldView-2/3/4、IKONOS、QuickBird、GaoFen-2)上进行零样本与单样本评估,采用全量与冻结微调策略。在多种网络架构(卷积神经网络、Transformer、Mamba)上实验表明,该预训练策略显著提升不同卫星传感器与成像条件下的泛化性能。预训练模型在零样本场景中表现优越,在单样本设置中仅需极少真实数据即展现强大适应能力。本工作为跨域全色锐化提供了实用方案,建立了遥感图像融合任务的新泛化基准,推动了基础模型通过先进训练策略在遥感领域的应用。
原文摘要 · Abstract (English)
Existing deep learning methods for remote sensing image fusion often suffer from poor generalization when applied to unseen datasets due to the limited availability of real training data and the domain gap between different satellite sensors. To address this challenge, we explore the potential of foundation models by proposing a novel pretraining strategy that leverages large-scale simulated datasets to learn robust spatial-spectral priors. Specifically, our approach first constructs diverse simulated datasets by applying various degradation operations (blur, noise, downsampling) and augmentations (bands generation, channel shuffling, high-pass filtering, color jittering, etc.) to natural images from ImageNet and remote sensing images from SkyScript. We then pretrain fusion models on these simulated data to learn generalizable spatial-spectral representations. The pretrained models are subsequently evaluated on six datasets (WorldView-2/3/4, IKONOS, QuickBird, GaoFen-2) using zero-shot and one-shot paradigms, with both full- and freeze-tuning approaches for fine-tuning. Extensive experiments on different network architectures including convolutional neural networks, Transformer, and Mamba demonstrate that our pretraining strategy significantly improves generalization performance across different satellite sensors and imaging conditions for various fusion models. The pretrained models achieve superior results in zero-shot scenarios and show remarkable adaptation capability with minimal real data in one-shot settings. Our work provides a practical solution for cross-domain pansharpening, establishes a new benchmark for generalization in remote sensing image fusion tasks, and paves the way for leveraging foundation models through advanced training strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。