arXiv:2504.00901cs.CV2025-04被引 33

综述十年深度学习在遥感时空融合中的进展与挑战。

A Decade of Deep Learning for Remote Sensing Spatiotemporal Fusion: Advances, Challenges, and Opportunities

  • 按模型架构分类,梳理了CNN、Transformer等方法的演进
  • 验证了Transformer在长时序依赖建模上的优势,扩散模型细节重建更优
  • 指出跨数据集泛化、计算效率等五大关键难题,展望基础模型机遇

遥感时空融合(STF)通过融合高时频低空间与高空间低时频影像,缓解时空分辨率权衡问题。本文首次系统综述过去十年深度学习在遥感STF中的发展。我们构建了包括卷积神经网络(CNN)、Transformer、生成对抗网络(GAN)、扩散模型和序列模型在内的架构分类体系,揭示深度学习在该任务中应用的显著增长。分析表明,基于CNN的方法主导空间特征提取,而Transformer在捕捉长程时间依赖方面表现更优;GAN与扩散模型在细节重建上能力突出,显著优于传统方法,在结构相似性和光谱保真度上表现优异。通过对七个基准数据集上十种代表性方法的综合实验,验证了上述发现并量化了不同方法间的性能权衡。研究识别出五大关键挑战:时空冲突、跨数据集泛化能力有限、大规模处理计算效率低、多源异构融合困难以及基准多样性不足。论文提出未来在基础模型、混合架构和自监督学习方面的机遇,可能突破现有局限并支持多模态应用。文中提及的具体模型、数据集等信息已整理至:https://github.com/yc-cui/Deep-Learning-Spatiotemporal-Fusion-Survey。

原文摘要 · Abstract (English)

Remote sensing spatiotemporal fusion (STF) addresses the fundamental trade-off between temporal and spatial resolution by combining high temporal-low spatial and high spatial-low temporal imagery. This paper presents the first comprehensive survey of deep learning advances in remote sensing STF over the past decade. We establish a systematic taxonomy of deep learning architectures including Convolutional Neural Networks (CNNs), Transformers, Generative Adversarial Networks (GANs), diffusion models, and sequence models, revealing significant growth in deep learning adoption for STF tasks. Our analysis reveals that CNN-based methods dominate spatial feature extraction, while Transformer architectures show superior performance in capturing long-range temporal dependencies. GAN and diffusion models demonstrate exceptional capability in detail reconstruction, substantially outperforming traditional methods in structural similarity and spectral fidelity. Through comprehensive experiments on seven benchmark datasets comparing ten representative methods, we validate these findings and quantify the performance trade-offs between different approaches. We identify five critical challenges: time-space conflicts, limited generalization across datasets, computational efficiency for large-scale processing, multi-source heterogeneous fusion, and insufficient benchmark diversity. The survey highlights promising opportunities in foundation models, hybrid architectures, and self-supervised learning approaches that could address current limitations and enable multimodal applications. The specific models, datasets, and other information mentioned in this article have been collected in: https://github.com/yc-cui/Deep-Learning-Spatiotemporal-Fusion-Survey.

遥感融合深度学习时空建模综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。