用注意力解耦与自适应门控提升高光谱图像分类精度
Hyperspectral Image Classification via Transformer-based Spectral-Spatial Attention Decoupling and Adaptive Gating
- 分离空间与光谱注意力,精准捕捉关键信息
- 在IN/UP/KSC数据集上超越主流方法
- 小样本高噪声下仍具强泛化能力
深度神经网络在高光谱图像分类中面临高维数据、地物稀疏分布和光谱冗余等挑战,常导致过拟合并限制泛化能力。本文提出新型网络STNet,其核心为空间-光谱Transformer模块的双重创新:首先显式解耦空间与光谱注意力,确保对高光谱图像关键信息的针对性捕捉;其次引入两个功能独立的门控机制,在注意力融合层(自适应注意力融合门控)与特征变换内部(GFFN)实现智能调控。该设计相比传统卷积神经网络具备更优的特征提取与融合能力,且在小样本与高噪声场景中降低过拟合风险。STNet在不增加网络深度或宽度的前提下增强模型表达能力,在IN、UP和KSC数据集上表现优于主流高光谱图像分类方法。
原文摘要 · Abstract (English)
Deep neural networks face several challenges in hyperspectral image classification, including high-dimensional data, sparse distribution of ground objects, and spectral redundancy, which often lead to classification overfitting and limited generalization capability. To more effectively extract and fuse spatial context with fine spectral information in hyperspectral image (HSI) classification, this paper proposes a novel network architecture called STNet. The core advantage of STNet stems from the dual innovative design of its Spatial-Spectral Transformer module: first, the fundamental explicit decoupling of spatial and spectral attention ensures targeted capture of key information in HSI; second, two functionally distinct gating mechanisms perform intelligent regulation at both the fusion level of attention flows (adaptive attention fusion gating) and the internal level of feature transformation (GFFN). This characteristic demonstrates superior feature extraction and fusion capabilities compared to traditional convolutional neural networks, while reducing overfitting risks in small-sample and high-noise scenarios. STNet enhances model representation capability without increasing network depth or width. The proposed method demonstrates superior performance on IN, UP, and KSC datasets, outperforming mainstream hyperspectral image classification approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。