arXiv:2602.01906cs.CVcs.AI2026-02

提出新型光谱注意力模型,提升高光谱图像分类精度与效率

DSXFormer: Dual-Pooling Spectral Squeeze-Expansion and Dynamic Context Attention Transformer for Hyperspectral Image Classification

  • 双池化光谱压缩扩张机制,增强光谱特征区分能力
  • 动态上下文注意力降低计算开销,同时捕捉局部谱空间关系
  • 在四个基准数据集上均超越现有方法,最高准确率达99.95%

高光谱图像分类(HSIC)因光谱维度高、谱-空关系复杂且标注样本有限而具有挑战性。尽管基于Transformer的模型在该任务中展现潜力,但现有方法常难以兼顾光谱判别力与计算效率。为此,本文提出DSXFormer:一种结合双池化光谱压缩扩张(DSX)块与动态上下文注意力(DCA)的新型Transformer架构。DSX块通过全局平均池化与最大池化的互补性,自适应重校准光谱通道,强化谱间依赖建模;DCA机制在窗口化Transformer中动态捕获局部谱-空关系,显著降低计算负担。二者协同实现谱重点与空间上下文表征的平衡。同时采用分块提取、嵌入与合并策略,支持多尺度特征学习。在萨利纳斯(SA)、印第安那松(IP)、帕维亚大学(PU)和肯尼迪航天中心(KSC)四个常用数据集上的实验表明,DSXFormer持续优于现有最优方法,分类准确率分别为99.95%、98.91%、99.85%和98.52%。

原文摘要 · Abstract (English)

Hyperspectral image classification (HSIC) is a challenging task due to high spectral dimensionality, complex spectral-spatial correlations, and limited labeled training samples. Although transformer-based models have shown strong potential for HSIC, existing approaches often struggle to achieve sufficient spectral discriminability while maintaining computational efficiency. To address these limitations, we propose a novel DSXFormer, a novel dual-pooling spectral squeeze-expansion transformer with Dynamic Context Attention for HSIC. The proposed DSXFormer introduces a Dual-Pooling Spectral Squeeze-Expansion (DSX) block, which exploits complementary global average and max pooling to adaptively recalibrate spectral feature channels, thereby enhancing spectral discriminability and inter-band dependency modeling. In addition, DSXFormer incorporates a Dynamic Context Attention (DCA) mechanism within a window-based transformer architecture to dynamically capture local spectral-spatial relationships while significantly reducing computational overhead. The joint integration of spectral dual-pooling squeeze-expansion and DCA enables DSXFormer to achieve an effective balance between spectral emphasis and spatial contextual representation. Furthermore, patch extraction, embedding, and patch merging strategies are employed to facilitate efficient multi-scale feature learning. Extensive experiments conducted on four widely used hyperspectral benchmark datasets, including Salinas (SA), Indian Pines (IP), Pavia University (PU), and Kennedy Space Center (KSC), demonstrate that DSXFormer consistently outperforms state-of-the-art methods, achieving classification accuracies of 99.95%, 98.91%, 99.85%, and 98.52%, respectively.

高光谱分类Transformer注意力机制双池化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。