GAIA用自监督学习从卫星图像中提取大气动态特征,提升气象预测能力。
GAIA: A Foundation Model for Operational Atmospheric Dynamics
- 融合MAE与DINO的混合自监督模型,无须标签训练。
- 在30%-95%数据缺失下仍能精准补全,对气旋和降雨预测性能更优。
- 适合气象建模、气候分析等领域的研究者使用。
我们提出GAIA(Geospatial Artificial Intelligence for Atmospheres),一种融合掩码自编码器(MAE)与自蒸馏(DINO)的混合自监督地理空间基础模型,用于从全球静止卫星影像中生成语义丰富的表示。模型在2001-2015年共15年的全球红外观测数据上预训练,学习到解耦的表征,捕捉大气动力学而非简单的昼夜模式,表现为分布式的主成分结构与时间一致性。在不同数据缺失率(30%-95%)下均具备稳健重建能力,真实缺失数据补全表现优于基线。迁移至下游任务时,GAIA持续优于仅使用MAE的基准:气旋河分割F1值达0.58(对比0.52),热带气旋检测的风暴级召回率提升至81%(对比75%),早期预警准确率提高至29%(对比17%),降水估计性能保持竞争力。分析表明,混合目标促使模型学习跨多个主成分的空间连贯、以对象为中心的特征,而非集中于重建的局部表示。本工作证明,结合互补的自监督目标可获得更具迁移性的大气建模表征。模型权重与代码已公开于 https://huggingface.co/bcg-usra-nasa-gaia/GAIA-v1。
原文摘要 · Abstract (English)
We introduce GAIA (Geospatial Artificial Intelligence for Atmospheres), a hybrid self-supervised geospatial foundation model that fuses Masked Autoencoders (MAE) with self-distillation with no labels (DINO) to generate semantically rich representations from global geostationary satellite imagery. Pre-trained on 15 years of globally-merged infrared observations (2001-2015), GAIA learns disentangled representations that capture atmospheric dynamics rather than trivial diurnal patterns, as evidenced by distributed principal component structure and temporal coherence analysis. We demonstrate robust reconstruction capabilities across varying data availability (30-95% masking), achieving superior gap-filling performance on real missing data patterns. When transferred to downstream tasks, GAIA consistently outperforms an MAE-only baseline: improving atmospheric river segmentation (F1: 0.58 vs 0.52), enhancing tropical cyclone detection (storm-level recall: 81% vs 75%, early detection: 29% vs 17%), and maintaining competitive precipitation estimation performance. Analysis reveals that GAIA's hybrid objectives encourage learning of spatially coherent, object-centric features distributed across multiple principal components rather than concentrated representations focused on reconstruction. This work demonstrates that combining complementary self-supervised objectives yields more transferable representations for diverse atmospheric modeling tasks. Model weights and code are available at: https://huggingface.co/bcg-usra-nasa-gaia/GAIA-v1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。