arXiv:2507.12967cs.CV2025-07

用RGB预训练模型提升光谱重建,精准捕捉传感器未记录的光谱信息。

RGB Pre-Training Enhanced Unobservable Feature Latent Diffusion Model for Spectral Reconstruction

  • 通过分离不可见光谱特征,在紧凑隐空间中联合建模光谱与空间结构。
  • 两阶段训练:先压缩不可见特征,再基于RGB图像引导生成光谱分布。
  • 在光谱重建和下游重光照任务中达到当前最优性能。

光谱重建(SR)是图像处理中的关键问题,旨在从对应RGB图像中恢复高光谱图像(HSIs)。其核心难点在于估计未被RGB成像传感器捕获的不可见特征,该特征蕴含重要光谱信息。解决方案在于基于RGB图像有效构建光谱-空间联合分布以补充不可见特征。由于HSI与对应RGB图像具有相似的空间结构,因此可利用RGB预训练模型中的丰富空间知识来学习光谱-空间联合分布。为此,我们将RGB预训练隐扩散模型(RGB-LDM)扩展为不可见特征隐扩散模型(ULDM)用于SR。由于RGB-LDM及其对应的空间自编码器(SpaAE)已具备出色的空间知识表达能力,ULDM可专注于建模光谱结构。此外,将不可见特征从HSI中分离,减少了冗余光谱信息,使ULDM能在紧凑隐空间中学习联合分布。具体提出两阶段流程:第一阶段使用光谱不可见特征自编码器(SpeUAE)提取并压缩不可见特征至与RGB空间对齐的3D流形;第二阶段由SpeUAE和SpaAE分别编码光谱与空间结构,最终获得基于对应RGB图像引导的编码不可见特征分布的ULDM。在SR及下游重光照任务上的实验结果表明,所提方法达到当前最优性能。

原文摘要 · Abstract (English)

Spectral reconstruction (SR) is a crucial problem in image processing that requires reconstructing hyperspectral images (HSIs) from the corresponding RGB images. A key difficulty in SR is estimating the unobservable feature, which encapsulates significant spectral information not captured by RGB imaging sensors. The solution lies in effectively constructing the spectral-spatial joint distribution conditioned on the RGB image to complement the unobservable feature. Since HSIs share a similar spatial structure with the corresponding RGB images, it is rational to capitalize on the rich spatial knowledge in RGB pre-trained models for spectral-spatial joint distribution learning. To this end, we extend the RGB pre-trained latent diffusion model (RGB-LDM) to an unobservable feature LDM (ULDM) for SR. As the RGB-LDM and its corresponding spatial autoencoder (SpaAE) already excel in spatial knowledge, the ULDM can focus on modeling spectral structure. Moreover, separating the unobservable feature from the HSI reduces the redundant spectral information and empowers the ULDM to learn the joint distribution in a compact latent space. Specifically, we propose a two-stage pipeline consisting of spectral structure representation learning and spectral-spatial joint distribution learning to transform the RGB-LDM into the ULDM. In the first stage, a spectral unobservable feature autoencoder (SpeUAE) is trained to extract and compress the unobservable feature into a 3D manifold aligned with RGB space. In the second stage, the spectral and spatial structures are sequentially encoded by the SpeUAE and the SpaAE, respectively. The ULDM is then acquired to model the distribution of the coded unobservable feature with guidance from the corresponding RGB images. Experimental results on SR and downstream relighting tasks demonstrate that our proposed method achieves state-of-the-art performance.

光谱重建扩散模型隐空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。