用轻量解码器让气象大模型预测未训练过的水文变量,提速50%还省内存。
Finetuning a Weather Foundation Model with Lightweight Decoders for Unseen Physical Processes
- 用浅层解码器在预训练模型隐空间上训练,预测新变量
- 比全模型微调快50%,省35%内存,精度仍高
- 适合计算资源有限的研究者,扩展性强
近期人工智能天气预报发展催生了所谓“基础模型”,通常需昂贵预训练和少量微调。但在自然科学中,理想的模型应编码底层物理变量间的有意义统计关系。本研究评估了当前最先进的Aurora基础模型在预测预训练时未包含的水文变量上的表现。提出一种轻量级方法:在预训练模型的隐表示上训练浅层解码器来预测新变量。作为基线,对比了全模型微调(可优化隐空间并引入新变量输入输出)。解码器方法训练时间减少50%,内存降低35%,在多种水文变量上保持高精度,同时保留模型自回归稳定性等优良特性。值得注意的是,解码器精度与新变量和预训练变量间的物理相关性正相关,表明Aurora隐空间捕捉了有意义的物理关系。因此我们认为,地球科学中基础模型的重要评价标准是能否无需全量微调即可扩展至新变量。这为计算资源有限的群体提供了新路径,推动基础模型在地球科学中的更广泛应用。
原文摘要 · Abstract (English)
Recent advances in AI weather forecasting have led to the emergence of so-called "foundation models", typically defined by expensive pretraining and minimal fine-tuning for downstream tasks. However, in the natural sciences, a desirable foundation model should also encode meaningful statistical relationships between the underlying physical variables. This study evaluates the performance of the state-of-the-art Aurora foundation model in predicting hydrological variables, which were not considered during pretraining. We introduce a lightweight approach using shallow decoders trained on the latent representations of the pretrained model to predict these new variables. As a baseline, we compare this to fine-tuning the full model, which allows further optimization of the latent space while incorporating new variables into both inputs and outputs. The decoder-based approach requires 50% less training time and 35% less memory, while achieving strong accuracy across various hydrological variables and preserving desirable properties of the foundation model, such as autoregressive stability. Notably, decoder accuracy depends on the physical correlation between the new variables and those used during pretraining, indicating that Aurora's latent space captures meaningful physical relationships. In this sense, we argue that an important quality metric for foundation models in Earth sciences is their ability to be extended to new variables without a full fine-tuning. This provides a new perspective for making foundation models more accessible to communities with limited computational resources, while supporting broader adoption in Earth sciences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。