arXiv:2502.09356cs.CV2025-02ICML被引 54

统一处理多种遥感数据,实现跨尺度精准识别。

Galileo: Learning Global & Local Features of Many Remote Sensing Modalities

  • 设计双分支对比损失,同时捕捉全局与局部特征。
  • 在11个基准上超越现有专用模型,提升多任务表现。
  • 适合遥感图像分析、灾害监测等需要多模态融合的场景。

我们提出一种高度多模态的Transformer模型——Galileo,用于表征多种遥感模态(包括多光谱光学、合成孔径雷达、高程、气象数据、伪标签等)在时空维度上的信息。这些输入对作物制图、洪水检测等任务至关重要。然而,由于模态差异大且目标尺度跨度极广(从1-2像素的小船到数千像素的冰川),学习共享表示极具挑战。为此,我们设计了一种新颖的自监督学习算法,通过掩码建模提取跨尺度特征。模型采用双对比损失:全局损失基于深层表示,局部损失基于浅层输入投影;掩码策略分别为结构化和非结构化。Galileo是一个通用型单一模型,在卫星图像和像素时间序列任务中,于11个基准上均优于当前最优专用模型,覆盖多个任务。

原文摘要 · Abstract (English)

We introduce a highly multimodal transformer to represent many remote sensing modalities - multispectral optical, synthetic aperture radar, elevation, weather, pseudo-labels, and more - across space and time. These inputs are useful for diverse remote sensing tasks, such as crop mapping and flood detection. However, learning shared representations of remote sensing data is challenging, given the diversity of relevant data modalities, and because objects of interest vary massively in scale, from small boats (1-2 pixels and fast) to glaciers (thousands of pixels and slow). We present a novel self-supervised learning algorithm that extracts multi-scale features across a flexible set of input modalities through masked modeling. Our dual global and local contrastive losses differ in their targets (deep representations vs. shallow input projections) and masking strategies (structured vs. not). Our Galileo is a single generalist model that outperforms SoTA specialist models for satellite images and pixel time series across eleven benchmarks and multiple tasks.

遥感多模态自监督时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。