用多光谱信息提升红外图像色彩还原,让夜视画面更真实清晰。
FSCM: Frequency-Enhanced Spatial-Spectral Coupled Mamba for Infrared Hyperspectral Image Colorization

- 通过频域增强与空间谱耦合建模,精准恢复红外图像细节
- 在复杂道路场景中实现高保真色彩与语义一致性,视觉质量领先
- 适合需要夜间感知与跨模态模型迁移的自动驾驶研究者
热红外成像具有抗光照变化和烟雾干扰的优势,对全天候感知至关重要。然而,缺乏自然色彩与精细纹理限制了目标识别、人眼判读及可见光模型的迁移。现有红外着色方法多依赖单波段图像,光谱信息不足易导致结构失真与语义混淆。尽管红外高光谱图像提供丰富光谱响应与材料信息,但现有单波段框架仍难以建模空间-光谱耦合关系与弱纹理细节。为此,本文提出FSCM——一种光谱信息引导的生成对抗网络框架。其核心为级联的频域增强型空间-光谱状态空间生成器,由多个FSB单元构成:状态空间建模捕捉全局空间-光谱依赖;频域增强模块(FEM)结合多级小波分解与傅里叶门控,恢复结构轮廓、方向高频细节与全局频率响应;双流混合门控模块(DGM)融合形变感知采样与稀疏注意力,强化有效局部结构并抑制背景干扰。此外,引入在线语义分割引导损失,提升复杂道路场景中的语义一致性。实验表明,FSCM在视觉质量与语义保真度上均优于现有方法。
原文摘要 · Abstract (English)
Thermal infrared imaging is robust to illumination variations and smoke interference, making it important for all-weather perception. However, the lack of natural color and fine texture limits target recognition, human visual interpretation, and the transfer of visible-light models. Existing infrared colorization methods mainly rely on single-band images, where insufficient spectral cues may lead to structural distortion and semantic confusion. Although infrared hyperspectral images provide rich spectral responses and material information, existing single-band frameworks remain limited in modeling spatial-spectral coupling and weak texture details. To address these issues, this paper presents FSCM, a spectral-information-guided GAN framework. Within FSCM, a frequency-enhanced spatial-spectral state-space generator composed of cascaded FSB units is constructed. Each FSB integrates three complementary components: state-space modeling captures global spatial-spectral dependencies; the frequency enhancement module (FEM) combines multi-level wavelet decomposition and Fourier gating to recover structural contours, directional high-frequency details, and global frequency responses; and the dual-stream hybrid gating module (DGM) integrates deformation-aware sampling with sparse attention to enhance effective local structures and suppress background interference. Additionally, an online semantic segmentation-guided loss is introduced to constrain the generated results, improving semantic consistency in complex road scenes. Experiments show that FSCM outperforms existing infrared colorization methods in visual quality and semantic fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。