arXiv:2509.10919cs.CVcs.LG2025-09被引 2

250万参数小模型,用元数据提升遥感图像预训练效果

Lightweight Metadata-Aware Mixture-of-Experts Masked Autoencoder for Earth Observation

  • 用稀疏专家路由+时空元数据编码,构建轻量级遥感自编码器
  • 在BigEarthNet上性能媲美大模型,线性探测准确率达83.7%
  • 无需显式元数据也能泛化,适合资源受限的遥感应用

近年来地球观测(EO)领域聚焦于大规模基础模型,但其计算开销大,限制了下游任务的可及性和复用性。本文探索紧凑架构作为构建小型通用EO模型的可行路径,提出仅含250万参数的元数据感知混合专家自编码器(MoE-MAE)。该模型结合稀疏专家路由与地理时间条件建模,将影像数据与经纬度、季节/每日周期编码一同输入。在BigEarthNet-Landsat数据集上进行预训练,并使用线性探针评估冻结编码器的嵌入表示。尽管模型规模小,仍能与更大模型竞争,表明元数据感知预训练提升了迁移能力和标签效率。为进一步评估泛化能力,在无显式元数据的EuroSAT-Landsat数据集上测试,结果仍优于数百兆参数模型。这表明紧凑的元数据感知MoE-MAE是迈向未来高效可扩展的地球观测基础模型的重要一步。

原文摘要 · Abstract (English)

Recent advances in Earth Observation have focused on large-scale foundation models. However, these models are computationally expensive, limiting their accessibility and reuse for downstream tasks. In this work, we investigate compact architectures as a practical pathway toward smaller general-purpose EO models. We propose a Metadata-aware Mixture-of-Experts Masked Autoencoder (MoE-MAE) with only 2.5M parameters. The model combines sparse expert routing with geo-temporal conditioning, incorporating imagery alongside latitude/longitude and seasonal/daily cyclic encodings. We pretrain the MoE-MAE on the BigEarthNet-Landsat dataset and evaluate embeddings from its frozen encoder using linear probes. Despite its small size, the model competes with much larger architectures, demonstrating that metadata-aware pretraining improves transfer and label efficiency. To further assess generalization, we evaluate on the EuroSAT-Landsat dataset, which lacks explicit metadata, and still observe competitive performance compared to models with hundreds of millions of parameters. These results suggest that compact, metadata-aware MoE-MAEs are an efficient and scalable step toward future EO foundation models.

遥感轻量化元数据自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。