arXiv:2409.09244cs.CV2024-09被引 13

提出新型光谱视觉变压器架构,提升遥感图像分类性能。

Investigation of Hierarchical Spectral Vision Transformer Architecture for Classification of Hyperspectral Imagery

  • 设计统一的分层光谱视觉变压器,融合卷积与自注意力模块。
  • 实验证明其在多种数据集上优于传统CNN,尤其抗干扰能力更强。
  • 适合遥感、高光谱图像分析研究人员参考使用。

过去三年,基于视觉变压器(Vision Transformer)的高光谱影像(HSI)分类在遥感数据分析中受到广泛关注。以往研究主要通过集成卷积神经网络(CNN)来增强局部特征提取能力,但视觉变压器在HSI分类中表现更优的理论依据仍不明确。为此,本文提出一种专为HSI分类设计的统一分层光谱视觉变压器架构。该架构在统一框架下整合多种混合模块:包括执行卷积操作的CNN-mixer,以及空间自注意力(SSA)和通道自注意力(CSA)两种改进型自注意力块;还有如SSA+CNN-mixer、CSA+CNN-mixer等混合模型,将卷积与自注意力结合。该设计可灵活构建多样化的视觉变压器模型。训练过程中,系统对比了经典CNN与视觉变压器模型,重点分析了扰动鲁棒性及海瑟矩阵最大特征值的分布。实验表明,视觉变压器的优势源于整体架构设计,而非单一多头自注意力(MSA)组件的贡献。

原文摘要 · Abstract (English)

In the past three years, there has been significant interest in hyperspectral imagery (HSI) classification using vision Transformers for analysis of remotely sensed data. Previous research predominantly focused on the empirical integration of convolutional neural networks (CNNs) to augment the network's capability to extract local feature information. Yet, the theoretical justification for vision Transformers out-performing CNN architectures in HSI classification remains a question. To address this issue, a unified hierarchical spectral vision Transformer architecture, specifically tailored for HSI classification, is investigated. In this streamlined yet effective vision Transformer architecture, multiple mixer modules are strategically integrated separately. These include the CNN-mixer, which executes convolution operations; the spatial self-attention (SSA)-mixer and channel self-attention (CSA)-mixer, both of which are adaptations of classical self-attention blocks; and hybrid models such as the SSA+CNN-mixer and CSA+CNN-mixer, which merge convolution with self-attention operations. This integration facilitates the development of a broad spectrum of vision Transformer-based models tailored for HSI classification. In terms of the training process, a comprehensive analysis is performed, contrasting classical CNN models and vision Transformer-based counterparts, with particular attention to disturbance robustness and the distribution of the largest eigenvalue of the Hessian. From the evaluations conducted on various mixer models rooted in the unified architecture, it is concluded that the unique strength of vision Transformers can be attributed to their overarching architecture, rather than being exclusively reliant on individual multi-head self-attention (MSA) components.

高光谱分类视觉变压器遥感图像特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。