arXiv:2510.27219cs.CV2025-10被引 4

提出可统一多传感器高光谱映射的新型基础模型,提升跨设备泛化能力。

SpecAware: A Spectral-Content Aware Foundation Model for Unifying Multi-Sensor Learning in Hyperspectral Remote Sensing Mapping

  • 通过融合传感器元属性与图像内容,动态生成每张图的条件输入。
  • 设计超网络驱动的双步矩阵分解,实现空间-光谱特征统一表征。
  • 在40万+高质光谱块数据上预训练,支持跨传感器联合学习。

高光谱成像(HSI)是精细土地利用与土地覆盖(LULC)制图的关键技术。然而,不同传感器间光谱通道的固有异质性长期制约了迁移学习或联合训练中的模型泛化。现有HSI基础模型虽在下游任务中表现良好,但普遍忽视传感器元属性与图像语义特征的关键引导作用,导致跨传感器联合学习适应性有限。为此,本文提出SpecAware,一种面向高光谱映射的多传感器统一学习基础模型。为支持该工作,我们构建了包含超过400,000个高质量光谱块的Hyper-400K数据集,覆盖多种机载AVIRIS传感器及两种数据处理层级(L1和L2)。SpecAware核心是一个超网络驱动的统一图像嵌入流程:首先设计元内容感知模块,融合传感器元属性与图像内容,为每个光谱样本生成独特条件输入;其次设计HyperEmbedding模块,通过样本条件超网络动态生成一对矩阵因子,实现通道级编码。该过程包括自适应空间模式提取与潜在语义特征投影,生成统一的高光谱令牌表示。因此,SpecAware可在多样场景与传感器间捕捉并解析空间-光谱特征,实现在统一多传感器联合预训练框架下的可适应性处理。

原文摘要 · Abstract (English)

Hyperspectral imaging (HSI) is a critical technique for fine-grained land-use and land-cover (LULC) mapping. However, the inherent heterogeneity of HSI data, particularly the variation in spectral channels across sensors, has long constrained the development of model generalization via transfer learning or joint training. Existing HSI foundation models show promise for different downstream tasks, but typically underutilize the critical guiding role of sensor meta-attributes and image semantic features, resulting in limited adaptability to cross-sensor joint learning. To address these issues, we propose SpecAware, which is a novel hyperspectral spectral-content aware foundation model for unifying multi-sensor learning for HSI mapping. To support this work, we constructed the Hyper-400K dataset, which is a new large-scale pre-training dataset with over 400\,k high-quality patches from diverse airborne AVIRIS sensors that cover two data processing levels (L1 and L2). The core of SpecAware is a hypernetwork-driven unified image embedding process for HSI data. Firstly, we designed a meta-content aware module to generate a unique conditional input for each HSI sample, tailored to each spectral band by fusing the sensor meta-attributes and its own image content. Secondly, we designed the HyperEmbedding module, where a sample-conditioned hypernetwork dynamically generates a pair of matrix factors for channel-wise encoding. This process implements two-step matrix factorization, consisting of adaptive spatial pattern extraction and latent semantic feature projection, yielding a unified hyperspectral token representation. Thus, SpecAware learns to capture and interpret spatial-spectral features across diverse scenes and sensors, enabling adaptive processing of variable spectral channels within a unified multi-sensor joint pre-training framework.

高光谱多传感器基础模型图像表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。