提出轻量高效多模态遥感模型,提升跨卫星泛化能力与计算效率。
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data
- 用低秩分解近似空间光谱注意力,降低计算开销。
- 在跨卫星任务中性能超越现有模型,且效率更高。
- 适合处理多通道、多模态的海量遥感数据研究者。
卫星多时相、多光谱遥感栅格数据蕴含丰富的时空上下文信息,具有广泛应用潜力。现有自监督学习方法在面对不断增加的通道与模态时,受限于可扩展性,存在灵活性差和计算效率低的问题。为此,本文提出低秩高效空间-光谱视觉变换器(LESS ViT),包含三项创新:(i) LESS注意力模块通过克罗内克积近似高维空间-光谱注意力;(ii) 连续位置-通道嵌入层保持空间-光谱块的连续性与物理特性;(iii) 感知场掩码通过限制注意力至邻近区域利用局部空间依赖性。为评估方法,构建了GFM-Bench基准。采用高光谱掩码自编码器框架预训练,结合位置与通道掩码策略。实验表明,该方法在多模态遥感基础模型中表现优异,尤其在跨卫星泛化任务上优于当前最佳模型,且计算效率更高。其灵活性与可扩展性为未来多模态遥感分析提供了新方向。
原文摘要 · Abstract (English)
Geospatial raster data, such as that collected by satellite-based imaging systems at different times and spectral bands, hold immense potential for enabling a wide range of high-impact applications. This potential stems from the rich information that is spatially and temporally contextualized across multiple channels and sensing modalities. Recent work has adapted existing self-supervised learning approaches for such geospatial data. However, they fall short of scalable model architectures, leading to inflexibility and computational inefficiencies when faced with an increasing number of channels and modalities. To address these limitations, we introduce Low-rank Efficient Spatial-Spectral Vision Transformer with three key innovations: i) the LESS Attention Block that approximates high-dimensional spatial-spectral attention through Kronecker's product of the low-dimensional spatial and spectral attention components; ii) the Continuous Positional-Channel Embedding Layer that preserves both the continuity and physical characteristics of each spatial-spectral patch; and iii) the Perception Field Mask that exploits local spatial dependencies by constraining attention to neighboring patches. To evaluate the proposed innovations, we construct GFM-Bench, which serves as a comprehensive benchmark for such geospatial raster data. We pretrain LESS ViT using a Hyperspectral Masked Autoencoder framework with integrated positional and channel masking strategies. Experimental results demonstrate that our proposed method achieves competitive performance against state-of-the-art multi-modal geospatial foundation models while outperforming them on cross-satellite generalization tasks with higher computational efficiency. The flexibility and extensibility of our framework make it a promising direction for future geospatial data analysis tasks that involve a wide range of modalities and channels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。