统一多模态遥感模型,提升地球观测任务泛化能力
SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing

- 用单一视觉变压器处理多种遥感数据,避免冗余训练
- 新自监督预训练策略适配遥感图像复杂语义分布
- 引入可学习模态提示和专家混合模块,适合多任务应用
多模态遥感基础模型(MM-RSFM)显著推动了城市规划、环境监测和自然灾害管理等地球观测任务的发展。然而,现有方法通常为每种数据模态单独训练主干网络,造成参数冗余与利用效率低。同时,主流预训练方法沿用自然图像的自监督学习(SSL)策略,未能充分考虑遥感图像中单图内复杂的语义分布特性。本文提出 SkySense V2,一个统一的多模态遥感基础模型,采用单一 Transformer 主干处理多种模态数据,并通过专为遥感数据设计的新式自监督学习策略进行预训练。特别地,该模型引入自适应补丁合并模块与可学习模态提示令牌,以应对不同分辨率及模态间特征多样性不足的问题。此外,还融合专家混合(MoE)模块进一步提升性能。在涵盖7项任务、16个数据集的全面评估中,SkySense V2 相较于前代模型平均提升1.8点,展现出卓越的泛化能力。
原文摘要 · Abstract (English)
The multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for each data modality, leading to redundancy and inefficient parameter utilization. Moreover, prevalent pre-training methods typically apply self-supervised learning (SSL) techniques from natural images without adequately accommodating the characteristics of remote sensing (RS) images, such as the complicated semantic distribution within a single RS image. In this work, we present SkySense V2, a unified MM-RSFM that employs a single transformer backbone to handle multiple modalities. This backbone is pre-trained with a novel SSL strategy tailored to the distinct traits of RS data. In particular, SkySense V2 incorporates an innovative adaptive patch merging module and learnable modality prompt tokens to address challenges related to varying resolutions and limited feature diversity across modalities. In additional, we incorporate the mixture of experts (MoE) module to further enhance the performance of the foundation model. SkySense V2 demonstrates impressive generalization abilities through an extensive evaluation involving 16 datasets over 7 tasks, outperforming SkySense by an average of 1.8 points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。