用预训练大模型提升高光谱目标跟踪精度,仅需少量训练即可达成优异效果。
Spectral-Enhanced Transformers: Leveraging Large-Scale Pretrained Models for Hyperspectral Object Tracking
- 设计可学习的时空谱融合模块,适配任意Transformer主干网络
- 跨模态训练使模型在不同传感器数据上表现更鲁棒
- 在少样本情况下仍保持高性能,适合数据稀缺场景
基于快照马赛克相机的高光谱目标跟踪正逐渐兴起,因其能同时提供丰富的光谱信息与空间数据,有助于更全面地理解物质属性。尽管基于Transformer的模型在学习特征表示方面已显著优于卷积神经网络(CNN),但训练大型Transformer需要大规模数据集和长时间训练,而高光谱领域缺乏充足数据,制约了其性能发挥。本文提出一种有效方法,将大型预训练Transformer基础模型迁移应用于高光谱目标跟踪。我们设计了一种自适应、可学习的时空谱令牌融合模块,可嵌入任意Transformer主干网络,以学习高光谱数据中的内在时空谱特征。此外,模型引入跨模态训练流程,促进在不同传感器模态采集的数据间进行有效学习。该机制可提取额外模态的互补知识,即使测试时未出现也可受益。实验表明,该方法仅需极少训练迭代即实现优异性能。
原文摘要 · Abstract (English)
Hyperspectral object tracking using snapshot mosaic cameras is emerging as it provides enhanced spectral information alongside spatial data, contributing to a more comprehensive understanding of material properties. Using transformers, which have consistently outperformed convolutional neural networks (CNNs) in learning better feature representations, would be expected to be effective for Hyperspectral object tracking. However, training large transformers necessitates extensive datasets and prolonged training periods. This is particularly critical for complex tasks like object tracking, and the scarcity of large datasets in the hyperspectral domain acts as a bottleneck in achieving the full potential of powerful transformer models. This paper proposes an effective methodology that adapts large pretrained transformer-based foundation models for hyperspectral object tracking. We propose an adaptive, learnable spatial-spectral token fusion module that can be extended to any transformer-based backbone for learning inherent spatial-spectral features in hyperspectral data. Furthermore, our model incorporates a cross-modality training pipeline that facilitates effective learning across hyperspectral datasets collected with different sensor modalities. This enables the extraction of complementary knowledge from additional modalities, whether or not they are present during testing. Our proposed model also achieves superior performance with minimal training iterations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。