arXiv:2502.19699cs.CV2025-02被引 5

用扩散模型+对比学习提升高光谱图像分类精度

Spatial-Spectral Diffusion Contrastive Representation Network for Hyperspectral Image Classification

  • 设计新型分阶段网络结构,融合空间与光谱自注意力机制
  • 引入对数绝对误差损失和对比学习,增强特征区分性
  • 自适应选择去噪时间步,提升特征融合与分类效果

尽管高效提取具有判别性的空谱特征对高光谱图像分类(HSIC)至关重要,但受空谱异质性和噪声影响,实现仍具挑战。本文提出基于去噪扩散概率模型(DDPM)与对比学习(CL)的空谱扩散对比表示网络(DiffCRN)。首先,设计新型分阶段架构,包含空间自注意力去噪模块(SSAD)与光谱组自注意力去噪模块(SGSAD),提升空谱特征学习效率。其次,引入对数绝对误差(LAE)损失与对比学习,增强无监督特征学习效果,提高实例级与类间可区分性。第三,提出基于像素级光谱角映射(SAM)的可学习时间步选择方法,实现自适应、自动的时间步筛选。最后,设计自适应加权相加模块(AWAM)与跨时间步空谱融合模块(CTSSFM),有效融合时序特征并完成分类。在四个主流高光谱数据集上的实验表明,DiffCRN优于经典骨干模型及当前先进的GAN、Transformer与预训练方法。代码与预训练模型将公开。

原文摘要 · Abstract (English)

Although efficient extraction of discriminative spatial-spectral features is critical for hyperspectral images classification (HSIC), it is difficult to achieve these features due to factors such as the spatial-spectral heterogeneity and noise effect. This paper presents a Spatial-Spectral Diffusion Contrastive Representation Network (DiffCRN), based on denoising diffusion probabilistic model (DDPM) combined with contrastive learning (CL) for HSIC, with the following characteristics. First,to improve spatial-spectral feature representation, instead of adopting the UNets-like structure which is widely used for DDPM, we design a novel staged architecture with spatial self-attention denoising module (SSAD) and spectral group self-attention denoising module (SGSAD) in DiffCRN with improved efficiency for spectral-spatial feature learning. Second, to improve unsupervised feature learning efficiency, we design new DDPM model with logarithmic absolute error (LAE) loss and CL that improve the loss function effectiveness and increase the instance-level and inter-class discriminability. Third, to improve feature selection, we design a learnable approach based on pixel-level spectral angle mapping (SAM) for the selection of time steps in the proposed DDPM model in an adaptive and automatic manner. Last, to improve feature integration and classification, we design an Adaptive weighted addition modul (AWAM) and Cross time step Spectral-Spatial Fusion Module (CTSSFM) to fuse time-step-wise features and perform classification. Experiments conducted on widely used four HSI datasets demonstrate the improved performance of the proposed DiffCRN over the classical backbone models and state-of-the-art GAN, transformer models and other pretrained methods. The source code and pre-trained model will be made available publicly.

高光谱图像扩散模型对比学习空谱融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。