arXiv:2504.18770cs.CVcs.AI2025-04被引 2

用注意力机制融合多源遥感图像,构建可解释的地球观测基础模型

PyViT-FUSE: A Foundation Model for Multi-Sensor Earth Observation Data

  • 通过注意力机制融合任意数量、不同分辨率的遥感波段
  • 采用新型金字塔结构的视觉变压器堆叠处理特征表示
  • 自监督训练后在下游任务中表现良好,且融合过程可可视化

我们提出 PyViT-FUSE,一种专为多模态遥感数据设计的基础模型,通过注意力机制将任意数量、混合分辨率的输入波段融合为统一表征。学习得到的图像块标记进一步经由带有新颖金字塔结构的视觉变压器堆叠处理。模型在全球采样数据集上以自监督方式训练,借鉴 SwAV 算法的核心思想。通过可视化注意力得分展示了融合机制的可解释性,并验证了模型在下游任务中的适用性。

原文摘要 · Abstract (English)

We propose PyViT-FUSE, a foundation model for earth observation data explicitly designed to handle multi-modal imagery by learning to fuse an arbitrary number of mixed-resolution input bands into a single representation through an attention mechanism. The learned patch tokens are further processed by a stack of vision transformers with a novel pyramidal structure. We train the model on a globally sampled dataset in a self-supervised manner, leveraging core concepts of the SwAV algorithm. We show the interpretability of the fusion mechanism by visualization of the attention scores and the models applicability to downstream tasks.

遥感多模态视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。