arXiv:2505.03220cs.CV2025-05被引 2

用空间与频域双重掩码预训练,提升高光谱图像识别效果

Dual-Domain Masked Image Modeling: A Self-Supervised Pretraining Strategy Using Spatial and Frequency Domain Masking for Hyperspectral Data

  • 在空间和频域同时随机遮蔽图像块与频谱成分
  • 在3个公开数据集上达顶尖分类性能,微调时收敛更快
  • 适合缺乏标注数据的高光谱分析任务

高光谱图像(HSI)蕴含丰富的光谱信息,可揭示材料关键属性,在多个领域具有广泛应用。然而,标注数据稀缺限制了深度学习尤其是基于Transformer模型的潜力。为此,本文提出空间-频率掩码图像建模(SFMIM),一种利用大量未标注数据的自监督预训练策略。方法将高光谱立方体按空间维度划分为不重叠的块,每块包含对应位置的完整光谱。在空间掩码中,随机遮蔽部分块,训练模型用可见块重建被遮区域;在频率掩码中,移除输入光谱的部分频域成分,并预测缺失频段。通过重建被掩码内容,基于Transformer的编码器学习到更深层次的光谱-空间关联。我们在三个公开高光谱分类基准上评估该方法,结果表明其达到当前最优性能,且微调阶段收敛迅速,验证了预训练策略的高效性。

原文摘要 · Abstract (English)

Hyperspectral images (HSIs) capture rich spectral signatures that reveal vital material properties, offering broad applicability across various domains. However, the scarcity of labeled HSI data limits the full potential of deep learning, especially for transformer-based architectures that require large-scale training. To address this constraint, we propose Spatial-Frequency Masked Image Modeling (SFMIM), a self-supervised pretraining strategy for hyperspectral data that utilizes the large portion of unlabeled data. Our method introduces a novel dual-domain masking mechanism that operates in both spatial and frequency domains. The input HSI cube is initially divided into non-overlapping patches along the spatial dimension, with each patch comprising the entire spectrum of its corresponding spatial location. In spatial masking, we randomly mask selected patches and train the model to reconstruct the masked inputs using the visible patches. Concurrently, in frequency masking, we remove portions of the frequency components of the input spectra and predict the missing frequencies. By learning to reconstruct these masked components, the transformer-based encoder captures higher-order spectral-spatial correlations. We evaluate our approach on three publicly available HSI classification benchmarks and demonstrate that it achieves state-of-the-art performance. Notably, our model shows rapid convergence during fine-tuning, highlighting the efficiency of our pretraining strategy.

高光谱图像自监督学习Transformer掩码建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。