首个地面高光谱预训练模型,解决传感器差异与数据少难题。
HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

- 通过自适应通道嵌入统一不同传感器的光谱输入
- 利用伪标签融合多源信息,在3个任务上显著提升性能
- 仅需微调头部即可跨设备通用,适合实际部署
高光谱成像可提供数百个窄波段的丰富时空谱信息,用于精确材料识别。然而,由于传感器光谱配置差异、标注有限且标签体系不一,以及现有数据集规模小、场景多样性不足,地面高光谱预训练骨干网络仍为空白。为此,我们提出首个地面高光谱预训练骨干模型 HyperVision。首先,采用通道自适应动态嵌入机制,将异构输入映射到统一标记空间;其次,构建无监督表示学习框架:针对标注稀疏与标签不一致问题,引入多源伪标签方法,融合 SAM2 的空间结构与 HyperFree 的细粒度光谱物质信息;为扩充场景多样性和弥补数据量不足,采用跨模态知识蒸馏,将预训练 RGB 视觉模型中的丰富语义表征迁移至本模型。在来自26个多样化地面数据集的1.5万张图像上预训练后,HyperVision展现出卓越泛化能力。仅需高效头部微调,无需调整骨干参数,在三种下游任务中均优于现有特定任务方法,实现高达16.3%的相对提升(超光谱语义分割 $ ext{Acc}_{ ext{M}}$),对象追踪 AUC 提升2.1%,显著物检测 MAE 降低35.5%。代码与预训练模型已开源。
原文摘要 · Abstract (English)
While hyperspectral imaging provides rich spatial-spectral information across hundreds of narrow wavelength bands for precise material identification, ground-based hyperspectral pre-trained backbones remain absent, constrained by varying spectral configurations across sensors, limited annotations and heterogeneous labeling schemes, and the limited scale and scene diversity of existing datasets. To address these challenges and enable universal perception, we propose HyperVision, the first ground-based hyperspectral pre-trained backbone. First, to handle varying spectral configurations, HyperVision adopts a channel-adaptive dynamic embedding mechanism to map heterogeneous inputs into a unified token space. Second, we develop an unsupervised representation learning framework. Specifically, to address limited annotations and heterogeneous labeling schemes, a multi-source pseudo-labeling method is introduced to fuse spatial structures from SAM2 and fine-grained spectral material information from HyperFree. Furthermore, to enrich scene diversity and compensate for limited dataset scale, a cross-modal knowledge distillation mechanism is utilized to transfer rich semantic representations from a pre-trained RGB vision model to our backbone. Pre-trained on a collection of 15k images from 26 diverse ground-based datasets, HyperVision demonstrates exceptional generalization. Requiring only efficient head-only adaptation without adjusting backbone parameters, it outperforms state-of-the-art task-specific methods across three downstream tasks under varying sensor configurations, yielding up to a 16.3% relative improvement in hyperspectral semantic segmentation $\mathrm{Acc}_{\mathrm{M}}$, a 2.1% relative gain in object tracking AUC, and a 35.5% reduction in salient object detection MAE. The source code and pre-trained models are available at https://github.com/lronkitty/HyperVision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。