arXiv:2512.22814cs.LGphysics.ao-ph2025-12被引 5

用1万年模拟数据训练出能一步预测长期天气的AI模型。

Long-Range Distillation: Distilling 10,000 Years of Simulated Climate into Long Timestep AI Weather Models

  • 用自回归教师模型生成超长时段气候数据,训练单步预测的学生模型。
  • 在真实世界中达到与欧洲中期天气预报中心相当的亚季节预报精度。
  • 首次证明合成数据可有效提升长周期天气预测能力,适合气象与气候研究者。

准确的长期天气预报仍是人工智能模型的重大挑战,既因自回归推演中误差累积,也因用于训练的再分析数据集对气候变率的慢模式样本有限。大多数AI天气模型为自回归式,需多次迭代才能达到亚季节至季节(S2S)或季节级预报,常引发不稳定与校准问题。长时距概率模型可一步生成长期预报,但仅用40年再分析数据训练易过拟合,表明需数个数量级更多数据。本文提出长程蒸馏方法,让长时距概率“学生”模型通过短时距自回归“教师”模型生成的超大规模合成数据进行训练。以深度学习地球系统模型(DLESyM)为教师,生成超过10,000年的模拟气候数据,训练出跨多时间尺度的蒸馏学生模型。在理想模型实验中,蒸馏模型表现优于气候学基准,接近其教师模型的性能,且将数百次自回归步骤简化为单步。在真实世界中,经ERA5微调后,其S2S预报技能媲美欧洲中期天气预报中心(ECMWF)集合预报。模型性能随合成数据量增加而提升,即使数据量远超ERA5,仍具可扩展性。这是首次证明人工合成数据可用于提升长周期天气预测能力。

原文摘要 · Abstract (English)

Accurate long-range weather forecasting remains a major challenge for AI models, both because errors accumulate over autoregressive rollouts and because reanalysis datasets used for training offer a limited sample of the slow modes of climate variability underpinning predictability. Most AI weather models are autoregressive, producing short lead forecasts that must be repeatedly applied to reach subseasonal-to-seasonal (S2S) or seasonal lead times, often resulting in instability and calibration issues. Long-timestep probabilistic models that generate long-range forecasts in a single step offer an attractive alternative, but training on the 40-year reanalysis record leads to overfitting, suggesting orders of magnitude more training data are required. We introduce long-range distillation, a method that trains a long-timestep probabilistic "student" model to forecast directly at long-range using a huge synthetic training dataset generated by a short-timestep autoregressive "teacher" model. Using the Deep Learning Earth System Model (DLESyM) as the teacher, we generate over 10,000 years of simulated climate to train distilled student models for forecasting across a range of timescales. In perfect-model experiments, the distilled models outperform climatology and approach the skill of their autoregressive teacher while replacing hundreds of autoregressive steps with a single timestep. In the real world, they achieve S2S forecast skill comparable to the ECMWF ensemble forecast after ERA5 fine-tuning. The skill of our distilled models scales with increasing synthetic training data, even when that data is orders of magnitude larger than ERA5. This represents the first demonstration that AI-generated synthetic training data can be used to scale long-range forecast skill.

天气预报长时距预测合成数据蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。