研究气象模型的缩放规律,发现扩大数据比增大模型更有效。
Scaling Laws of Global Weather Models
- 通过分析模型规模、数据量和算力的关系,发现数据增长10倍可降损3.2倍。
- 在固定算力下,增加数据量比增大模型能带来更大性能提升。
- 气象模型应优先设计更宽架构,而非更深结构,适合气候预测研究者。
数据驱动模型正在革新天气预报。为优化训练效率与模型性能,本文分析了该领域内的经验缩放规律。研究考察了模型性能(验证损失)与三个关键因素——模型规模(N)、数据集规模(D)和算力预算(C)之间的关系。在多种模型中,Aurora表现出最强的数据缩放能力:训练数据增加10倍,验证损失最高可降低3.2倍。GraphCast具有最高的参数效率,但存在硬件利用率低的问题。算力最优分析表明,在固定算力预算下,分配资源用于更大规模的训练数据,比增加模型规模能带来更大的性能增益。此外,我们分析了模型结构,发现其缩放行为与语言模型截然不同:气象模型始终更偏好增加宽度而非深度。这些发现表明,未来气象模型应优先采用更宽的架构和更大的有效训练数据集以最大化预测性能。
原文摘要 · Abstract (English)
Data-driven models are revolutionizing weather forecasting. To optimize training efficiency and model performance, this paper analyzes empirical scaling laws within this domain. We investigate the relationship between model performance (validation loss) and three key factors: model size ($N$), dataset size ($D$), and compute budget ($C$). Across a range of models, we find that Aurora exhibits the strongest data-scaling behavior: increasing the training dataset by 10x reduces validation loss by up to 3.2x. GraphCast demonstrates the highest parameter efficiency, yet suffers from limited hardware utilization. Our compute-optimal analysis indicates that, under fixed compute budgets, allocating resources to more total training data yields greater performance gains than increasing model size. Furthermore, we analyze model shape and uncover scaling behaviors that differ fundamentally from those observed in language models: weather forecasting models consistently favor increased width over depth. These findings suggest that future weather models should prioritize wider architectures and larger effective training datasets to maximize predictive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。