让神经网络更准确地做气象数据同化,提升预报稳定性。
Jacobian-Enforced Neural Networks (JENN) for Improved Data Assimilation Consistency in Dynamical Models
- 通过强制神经网络学习真实动力系统的雅可比关系,增强其一致性。
- 在洛伦兹96模型上,显著降低线性近似与伴随模型的噪声。
- 无需重训练或改结构,可直接适配现有气象大模型。
基于机器学习的气象模型在预报方面表现优异,但在数据同化(DA)任务中仍不如传统数值天气预报(NWP)模型。本文提出雅可比强制神经网络(JENN)框架,旨在提升神经网络模拟动力系统时的数据同化一致性。以洛伦兹96模型为例,该方法通过显式强制雅可比关系来改善神经网络在DA中的适用性。网络结构包含40个输入神经元、两个各含256个双曲正切激活单元的隐藏层,以及40个无激活函数的输出层。采用两阶段训练:第一阶段使用标准预测-标签对建立基础预报能力;第二阶段引入定制损失函数,结合状态值预测的均方根误差(RMSE)与切线线性(TL)和伴随(AD)模拟结果的额外RMSE项,权重平衡预报精度与雅可比敏感性。为保证一致性,第二阶段使用物理模型计算的额外TL/AD输入-标签对进行训练。该方法无需从头训练或修改网络结构,可直接应用于GraphCast、NeuralGCM、Pangu或FuXi等预训练模型,实现对DA任务的最小重构适配。实验表明,JENN在保持非线性预报性能的同时,显著降低了TL、AD分量及整体雅可比矩阵的噪声。
原文摘要 · Abstract (English)
Machine learning-based weather models have shown great promise in producing accurate forecasts but have struggled when applied to data assimilation tasks, unlike traditional numerical weather prediction (NWP) models. This study introduces the Jacobian-Enforced Neural Network (JENN) framework, designed to enhance DA consistency in neural network (NN)-emulated dynamical systems. Using the Lorenz 96 model as an example, the approach demonstrates improved applicability of NNs in DA through explicit enforcement of Jacobian relationships. The NN architecture includes an input layer of 40 neurons, two hidden layers with 256 units each employing hyperbolic tangent activation functions, and an output layer of 40 neurons without activation. The JENN framework employs a two-step training process: an initial phase using standard prediction-label pairs to establish baseline forecast capability, followed by a secondary phase incorporating a customized loss function to enforce accurate Jacobian relationships. This loss function combines root mean square error (RMSE) between predicted and true state values with additional RMSE terms for tangent linear (TL) and adjoint (AD) emulation results, weighted to balance forecast accuracy and Jacobian sensitivity. To ensure consistency, the secondary training phase uses additional pairs of TL/AD inputs and labels calculated from the physical models. Notably, this approach does not require starting from scratch or structural modifications to the NN, making it readily applicable to pretrained models such as GraphCast, NeuralGCM, Pangu, or FuXi, facilitating their adaptation for DA tasks with minimal reconfiguration. Experimental results demonstrate that the JENN framework preserves nonlinear forecast performance while significantly reducing noise in the TL and AD components, as well as in the overall Jacobian matrix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。