用物理模型提升双相机压缩感知高光谱成像质量
PCMamba: Physics-Informed Cross-Modal State Space Model for Dual-Camera Compressive Hyperspectral Imaging
- 将热成像物理过程融入Mamba架构,实现轻量化建模
- 分离温度、发射率和纹理三类物理属性,提升重建精度
- 适合做高光谱成像、红外传感与物理信息融合研究者
全色(PAN)辅助的双相机压缩高光谱成像(DCCHI)是快照式高光谱成像的关键技术。现有方法主要显式地从2D压缩测量中提取光谱信息,从PAN图像中提取空间信息,导致高光谱图像(HSI)重建存在瓶颈。温度、发射率及物体间多次反射等物理因素在传感器获取热高光谱信号过程中起关键作用。受此启发,我们探索物理属性间的相互关系,为HSI重建提供更深层次理论支持。本文提出物理感知跨模态状态空间模型(PCMamba),将高光谱成像的前向物理过程嵌入Mamba的线性复杂度结构中,实现轻量级且高质量的HSI重建。具体地,分析热高光谱信号成像过程,使网络能解耦温度、发射率和纹理三个关键物理属性。通过充分挖掘2D测量和PAN图像中的潜在信息,基于物理驱动的合成过程重建高光谱图像。此外,设计了跨模态扫描Mamba块(CSMB),通过交叉扫描主干特征与PAN特征,引入像素级跨模态交互与位置归纳偏置。在真实与模拟数据集上的大量实验表明,本方法在定量与定性指标上均显著优于当前最优方法。
原文摘要 · Abstract (English)
Panchromatic (PAN) -assisted Dual-Camera Compressive Hyperspectral Imaging (DCCHI) is a key technology in snapshot hyperspectral imaging. Existing research primarily focuses on exploring spectral information from 2D compressive measurements and spatial information from PAN images in an explicit manner, leading to a bottleneck in HSI reconstruction. Various physical factors, such as temperature, emissivity, and multiple reflections between objects, play a critical role in the process of a sensor acquiring hyperspectral thermal signals. Inspired by this, we attempt to investigate the interrelationships between physical properties to provide deeper theoretical insights for HSI reconstruction. In this paper, we propose a Physics-Informed Cross-Modal State Space Model Network (PCMamba) for DCCHI, which incorporates the forward physical imaging process of HSI into the linear complexity of Mamba to facilitate lightweight and high-quality HSI reconstruction. Specifically, we analyze the imaging process of hyperspectral thermal signals to enable the network to disentangle the three key physical properties-temperature, emissivity, and texture. By fully exploiting the potential information embedded in 2D measurements and PAN images, the HSIs are reconstructed through a physics-driven synthesis process. Furthermore, we design a Cross-Modal Scanning Mamba Block (CSMB) that introduces inter-modal pixel-wise interaction with positional inductive bias by cross-scanning the backbone features and PAN features. Extensive experiments conducted on both real and simulated datasets demonstrate that our method significantly outperforms SOTA methods in both quantitative and qualitative metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。