arXiv:2409.02675eess.IVcs.CV2024-09IJCV被引 4

用深度展开网络融合卫星影像,提升分辨率与细节保真度。

Multi-Head Attention Residual Unfolded Network for Model-Based Pansharpening

  • 基于变分模型展开优化流程,结合多头注意力残差网络
  • 在三个数据集上实现优于现有方法的融合效果
  • 适合遥感图像处理、高光谱成像领域研究者使用

全色锐化与超色锐化的目标是将高分辨率全色(PAN)图像与低分辨率多光谱(MS)或超光谱(HS)图像精确融合。展开式融合方法将深度学习的强大表征能力与模型驱动方法的鲁棒性结合,通过将能量最小化的优化步骤展开为深度学习框架,形成高效且可解释的结构。本文提出一种基于模型的深度展开卫星图像融合方法。其核心为变分公式,包含经典的MS/HS观测模型、基于PAN图像的高频注入约束以及任意凸先验。在展开阶段,引入利用PAN图像编码几何信息的上采样与下采样层,基于残差网络实现。主干为多头注意力残差网络(MARNet),替代优化中的邻近算子,通过多头注意力与残差学习结合,利用基于块的非局部算子挖掘图像自相似性。此外,设计基于MARNet的后处理模块进一步提升融合质量。在PRISMA、Quickbird和WorldView2数据集上的实验表明,该方法性能优越,具备跨传感器配置及不同空间/光谱分辨率的泛化能力。源码将在https://github.com/TAMI-UIB/MARNet发布。

原文摘要 · Abstract (English)

The objective of pansharpening and hypersharpening is to accurately combine a high-resolution panchromatic (PAN) image with a low-resolution multispectral (MS) or hyperspectral (HS) image, respectively. Unfolding fusion methods integrate the powerful representation capabilities of deep learning with the robustness of model-based approaches. These techniques involve unrolling the steps of the optimization scheme derived from the minimization of an energy into a deep learning framework, resulting in efficient and highly interpretable architectures. In this paper, we propose a model-based deep unfolded method for satellite image fusion. Our approach is based on a variational formulation that incorporates the classic observation model for MS/HS data, a high-frequency injection constraint based on the PAN image, and an arbitrary convex prior. For the unfolding stage, we introduce upsampling and downsampling layers that use geometric information encoded in the PAN image through residual networks. The backbone of our method is a multi-head attention residual network (MARNet), which replaces the proximity operator in the optimization scheme and combines multiple head attentions with residual learning to exploit image self-similarities via nonlocal operators defined in terms of patches. Additionally, we incorporate a post-processing module based on the MARNet architecture to further enhance the quality of the fused images. Experimental results on PRISMA, Quickbird, and WorldView2 datasets demonstrate the superior performance of our method and its ability to generalize across different sensor configurations and varying spatial and spectral resolutions. The source code will be available at https://github.com/TAMI-UIB/MARNet.

图像融合遥感注意力机制深度展开

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。