arXiv:2605.05688cs.CV2026-05

用扩散模型实现高保真多光谱图像重建,仅需5步即可完成。

R2H-Diff: Guided Spectral Diffusion Model for RGB-to-Hyperspectral Reconstruction

  • 基于条件迭代优化,通过RGB引导逐步还原光谱信息。
  • 在NTIRE2022上达35.37 dB PSNR,参数仅0.58M,FLOPs为12.25G。
  • 创新设计模块提升光谱一致性,适合低资源场景下的高质量重建。

RGB到多光谱图像重建是一个高度病态的逆问题,因为多个合理的光谱分布可能对应相同的RGB观测值。现有回归方法通常学习确定性映射,难以建模重建不确定性,常导致光谱响应过度平滑。尽管扩散模型具备强大的分布建模能力,但其直接应用于多光谱重建仍面临高光谱维度、强波段相关性及严格的光谱保真度要求等挑战。为此,我们提出R2H-Diff,一种专用于RGB到多光谱图像重建的高效扩散框架。具体地,将光谱恢复建模为条件迭代精炼过程,在RGB引导下实现渐进式重建。设计了基于RGB条件特征融合的引导光谱精炼模块,以及用于高效空间-光谱依赖建模的多光谱自适应转置注意力模块。此外,采用无归一化去噪主干网络以保持光谱振幅一致性,并引入任务适配的线性噪声调度,在仅5步去噪的情况下实现高质量重建。在NTIRE2022、CAVE和哈佛数据集上的大量实验表明,R2H-Diff在重建质量与计算效率之间取得了良好平衡。值得注意的是,在NTIRE2022上,R2H-Diff以0.58M参数和12.25G FLOPs达到35.37 dB PSNR,是所评估方法中模型复杂度最低的,同时保持了强重建保真度。

原文摘要 · Abstract (English)

RGB-to-hyperspectral image reconstruction is a highly ill-posed inverse problem, since multiple plausible spectral distributions may correspond to the same RGB observation. Existing regression-based methods usually learn a deterministic mapping, which limits their ability to model reconstruction uncertainty and often leads to over-smoothed spectral responses. Although diffusion models provide strong distribution modeling capability, their direct application to hyperspectral reconstruction remains challenging due to the high spectral dimensionality, strong inter-band correlations, and strict requirement for spectral fidelity. To this end, we propose R2H-Diff, an efficient diffusion-based framework tailored for RGB-to-HSI reconstruction. Specifically, R2H-Diff formulates spectral recovery as a conditional iterative refinement process, enabling progressive reconstruction under RGB guidance. We proposed a Guided Spectral Refinement Module for RGB-conditioned feature fusion and a Hyperspectral-Adaptive Transposed Attention module for efficient spatial--spectral dependency modeling. Furthermore, a normalization-free denoising backbone is adopted to preserve spectral amplitude consistency, while a task-adapted linear noise schedule enables high-quality reconstruction with only five denoising steps. Extensive experiments on NTIRE2022, CAVE, and Harvard demonstrate that R2H-Diff achieves a favorable balance between reconstruction quality and computational efficiency. Notably, on NTIRE2022, R2H-Diff obtains 35.37 dB PSNR with a sub-million-parameter model of 0.58M parameters and 12.25G FLOPs, achieving the lowest model complexity among the evaluated methods while maintaining strong reconstruction fidelity.

多光谱重建扩散模型图像生成低参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。