用双模态提示引导扩散模型,实现无训练高光谱图像超分辨率。
Dual Modality Prompted Diffusion Priors for Zero Shot Hyperspectral Pansharpening

- 将低分辨率光谱与全色图像转为提示词,注入扩散模型中间层
- 在三种数据集上均超越现有方法,全分辨率下保真度最高
- 适合需要零样本、高保真的遥感图像处理场景
高光谱图像超分辨旨在从全色(PAN)和低分辨率高光谱(LRHS)图像重建高分辨率高光谱(HRHS)图像,同时保持空间细节与光谱保真度。现有基于扩散的方法通过生成低维表示并映射到高光谱域来利用预训练图像先验,但通常仅以外部重建目标引入观测图像,限制了其与扩散先验的直接交互。为此,本文提出双模态图像提示扩散模型(DIDM),分别将低分辨率高光谱与全色图像编码为光谱与空间提示词,通过交叉注意力注入冻结的遥感扩散模型中间特征,使互补的光谱与空间信息直接引导扩散特征演化。此外,引入全色引导加权像素感知总变差正则化器,结合低分辨率高光谱退化保真度与全色响应保真度,采用梯度自适应结构正则化,在保留结构不连续性的同时抑制均匀区域中的伪影。在Pavia、Chikusei、Houston数据集上,于降分辨率协议下的大量实验表明,DIDM在所有评估指标上表现最佳;在全分辨率评估(FR1)中,其获得最高HQNR。结果表明,内部双模态提示与全色引导结构正则化能有效平衡空间细节增强与光谱保真。
原文摘要 · Abstract (English)
Hyperspectral pansharpening aims to reconstruct a high resolution hyperspectral (HRHS) image from a panchromatic (PAN) image and a low resolution hyperspectral (LRHS) image while preserving both spatial details and spectral fidelity. Recent diffusion based methods exploit pretrained image priors by generating a low dimensional representation and subsequently mapping it to the hyperspectral domain. However, the observed panchromatic and hyperspectral images are typically imposed only through external reconstruction objectives, limiting their direct interaction with the diffusion prior. To address this issue, we propose dual-modality image-prompted diffusion model (DIDM) for zero shot hyperspectral pansharpening. DIDM encodes the low resolution hyperspectral and panchromatic observations into spectral and spatial prompt tokens, respectively, and injects them into intermediate features of a frozen remote sensing diffusion model through cross attention, allowing complementary spectral and spatial information to directly guide diffusion feature evolution. In addition, we introduce a panchromatic guided weighted pixel aware total variation regularizer that combines low resolution hyperspectral degradation fidelity and panchromatic response fidelity with gradient adaptive structural regularization, thereby preserving structural discontinuities while suppressing spurious variations in homogeneous regions. Extensive experiments on Pavia, Chikusei, and Houston under reduced resolution protocols show that DIDM achieves the best performance across all evaluated metrics, while full resolution evaluation on FR1 yields the highest HQNR among the compared methods. These results demonstrate that internal dual modality prompting and panchromatic guided structural regularization provide an effective balance between spatial detail enhancement and spectral preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。