arXiv:2605.21075cs.CVcs.LG2026-05被引 2

首个融合高光谱与多源遥感的预训练模型,提升地球观测理解能力。

SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining

论文配图:SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining
图 1 · 摘自论文原文
  • 设计分层架构,统一处理高光谱与多源遥感数据
  • 在200万全球点上预训练,覆盖40TB多模态数据
  • 适用于高光谱分析与多源遥感融合任务的研究者

地球观测基础模型(FM)正逐步采用多传感器数据,涵盖多光谱影像(MSI)、合成孔径雷达(SAR)及衍生地理空间图层,但高光谱影像(HSI)仍被忽视。现有高光谱模型仅基于HSI训练,缺乏与同位置其他地球观测传感器的联合预训练与融合研究。本文提出SpectralEarth-FM,一种针对异构光谱维度的多传感器地球观测分层变压器架构。该架构结合高光谱令牌化、传感器专用编码器、跨传感器融合模块和共享分层编码器,实现对高光谱与低通道观测数据的联合处理。为预训练此模型,我们构建SpectralEarth-MM数据集,将三颗星载高光谱传感器(EnMAP、EMIT、DESI)的数据与哨兵-2、陆地卫星-8/9光学影像、陆地表面温度(LST)及哨兵-1 SAR在共用地理区域上对齐,包含约200万全球分布地点、2500万个地理参考图像块和超过40TB数据。预训练采用类似JEPA的联合嵌入预测目标,匹配同一位置的全局视图与单传感器局部视图表示。我们在高光谱下游任务和标准地球观测基准测试中评估SpectralEarth-FM,遵循PANGAEA协议,结果在两种设置下均达到当前最优性能。

原文摘要 · Abstract (English)

Earth observation (EO) foundation models (FMs) are increasingly trained on multisensor data, spanning multispectral imagery (MSI), synthetic aperture radar (SAR), and derived geospatial layers, but hyperspectral imagery (HSI) remains underrepresented. Conversely, existing hyperspectral FMs are trained on HSI alone, leaving joint pretraining and fusion of HSI with co-located EO sensors unexplored. We introduce SpectralEarth-FM, a hierarchical transformer for multisensor EO input with heterogeneous spectral dimensionality. The architecture combines spectral tokenization for hyperspectral inputs, sensor-specific encoders, a cross-sensor fusion module, and a shared hierarchical encoder, enabling joint processing of HSI and lower-channel observations. To pretrain SpectralEarth-FM, we curate SpectralEarth-MM, a dataset that co-locates HSI from three spaceborne sensors (EnMAP, EMIT, DESIS) with Sentinel-2, Landsat-8/9 optical imagery, Landsat land surface temperature (LST), and Sentinel-1 SAR, over common geographic footprints. It comprises approximately 2M globally distributed locations, 25M georeferenced patches, and over 40TB of data. Pretraining uses a Joint-Embedding Predictive Architecture (JEPA)-style objective that matches representations between global views and single-sensor local views from the same location. We evaluate SpectralEarth-FM on hyperspectral downstream tasks and standard EO benchmarks following the PANGAEA protocol, achieving state-of-the-art results across both evaluation settings.

遥感多模态高光谱预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。