arXiv:2603.01174cs.CV2026-03被引 1

混合Mamba与Transformer,用图文提示提升高光谱图像分类精度

VP-Hype: A Hybrid Mamba-Transformer Framework with Visual-Textual Prompting for Hyperspectral Image Classification

  • 用混合Mamba-Transformer结构替代传统注意力,线性提升计算效率
  • 仅2%标注样本下,萨利纳斯数据集达到99.69%准确率,长口数据集99.45%
  • 适合小样本高光谱遥感分类任务,尤其关注计算效率与标签稀缺场景

高光谱图像(HSI)的精确分类常受限于高维光谱数据与极少量标注样本之间的矛盾。尽管如LoLA-SpecViT等层次化模型展示了局部窗口注意力和参数高效微调的优势,但标准Transformer的二次复杂度仍阻碍其扩展。本文提出VP-Hype框架,通过将状态空间模型(SSMs)的线性时间效率与Transformer的关系建模能力相结合,构建新型混合架构重新思考HSI分类。基于稳健的3D-CNN光谱前端,VP-Hype以混合Mamba-Transformer主干替换传统注意力模块,以显著降低计算开销的同时捕获长程依赖。此外,通过引入双模态视觉与文本提示,为特征提取过程提供上下文感知引导,缓解标签稀缺问题。实验表明,VP-Hype在低数据条件下建立新基准:在仅2%训练样本下,萨利纳斯数据集整体准确率达99.69%,长口数据集达99.45%。结果表明,混合序列建模与多模态提示的融合为高性能、样本高效的遥感分类提供了可靠路径。

原文摘要 · Abstract (English)

Accurate classification of hyperspectral imagery (HSI) is often frustrated by the tension between high-dimensional spectral data and the extreme scarcity of labeled training samples. While hierarchical models like LoLA-SpecViT have demonstrated the power of local windowed attention and parameter-efficient fine-tuning, the quadratic complexity of standard Transformers remains a barrier to scaling. We introduce VP-Hype, a framework that rethinks HSI classification by unifying the linear-time efficiency of State-Space Models (SSMs) with the relational modeling of Transformers in a novel hybrid architecture. Building on a robust 3D-CNN spectral front-end, VP-Hype replaces conventional attention blocks with a Hybrid Mamba-Transformer backbone to capture long-range dependencies with significantly reduced computational overhead. Furthermore, we address the label-scarcity problem by integrating dual-modal Visual and Textual Prompts that provide context-aware guidance for the feature extraction process. Our experimental evaluation demonstrates that VP-Hype establishes a new state of the art in low-data regimes. Specifically, with a training sample distribution of only 2\%, the model achieves Overall Accuracy (OA) of 99.69\% on the Salinas dataset and 99.45\% on the Longkou dataset. These results suggest that the convergence of hybrid sequence modeling and multi-modal prompting provides a robust path forward for high-performance, sample-efficient remote sensing.

高光谱分类混合模型小样本学习图文提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。