arXiv:2505.14951cs.CV2025-05被引 5

用多模态自编码器预训练地球观测模型,提升下游任务泛化能力。

MultiMAE Meets Earth Observation: Pre-training Multi-modal Multi-task Masked Autoencoders for Earth Observation Tasks

  • 采用多模态多任务掩码自编码器,统一重建光谱、高程和分割数据。
  • 在多个地球观测数据集上分类与分割任务性能超越现有方法。
  • 无需为每种模态单独预训练,适配不同输入组合,灵活性强。

地球观测(EO)中的多模态数据为提升深度学习模型的迁移学习能力提供了巨大机遇。与以往忽略多模态数据的研究不同,近期方法已开始引入多模态信息,实现更有效的预训练策略。然而,现有方法在将知识迁移到下游任务时仍面临挑战,尤其当下游数据结构与预训练阶段不一致时。本文提出一种更灵活的多模态多任务预训练策略,采用多模态多任务掩码自编码器(MultiMAE),通过重建包括光谱、高程和分割在内的多种输入模态进行预训练。预训练模型展现出强大的迁移能力,在多个地球观测数据集上的分类与分割任务中表现优于当前最优方法。该方法具有显著灵活性,可处理多样化的输入配置,无需为每种模态构建独立的预训练模型。代码将公开于:https://github.com/josesosajs/multimae-meets-eo。

原文摘要 · Abstract (English)

Multi-modal data in Earth Observation (EO) presents a huge opportunity for improving transfer learning capabilities when pre-training deep learning models. Unlike prior work that often overlooks multi-modal EO data, recent methods have started to include it, resulting in more effective pre-training strategies. However, existing approaches commonly face challenges in effectively transferring learning to downstream tasks where the structure of available data differs from that used during pre-training. This paper addresses this limitation by exploring a more flexible multi-modal, multi-task pre-training strategy for EO data. Specifically, we adopt a Multi-modal Multi-task Masked Autoencoder (MultiMAE) that we pre-train by reconstructing diverse input modalities, including spectral, elevation, and segmentation data. The pre-trained model demonstrates robust transfer learning capabilities, outperforming state-of-the-art methods on various EO datasets for classification and segmentation tasks. Our approach exhibits significant flexibility, handling diverse input configurations without requiring modality-specific pre-trained models. Code will be available at: https://github.com/josesosajs/multimae-meets-eo.

地球观测多模态自编码器迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。