用机器学习提升海洋生物地球化学数据同化,让模型更准地预测生态变化。
Hybrid machine learning data assimilation for marine biogeochemistry
- 用机器学习学习观测与未观测变量间的统计关系,改进数据同化
- 新方法使未直接观测变量更新效果显著优于传统单变量方案
- 模型具备一定跨区域迁移能力,适合未来扩展至三维业务系统
海洋生物地球化学模型对气候与人类活动影响的预测至关重要。数据同化(DA)通过将模型与真实观测对齐来提升模型性能,但受制于模型复杂性、强非线性及观测稀疏不确定,现有方法难以有效更新未观测变量;而基于集合的方法对高复杂度模型计算成本过高。本研究展示如何利用机器学习(ML)改善海洋生物地球化学数据同化:在西北欧洲海区1维原型系统中,集成机器学习驱动的平衡机制,用于从自由运行集合中学习状态依赖的相关性,并以“端到端”方式由集合卡尔曼滤波生成分析增量。结果表明,相比操作中常用的单变量方案,该方法显著提升了未观测变量的更新效果;且机器学习模型展现出中等程度的跨区域迁移能力,是迈向三维业务系统的关键一步。结论认为,机器学习为突破当前海洋生物地球化学数据同化计算瓶颈提供了清晰路径,未来需重点优化迁移性、训练数据采样策略并评估大规模应用的可扩展性。
原文摘要 · Abstract (English)
Marine biogeochemistry models are critical for forecasting, as well as estimating ecosystem responses to climate change and human activities. Data assimilation (DA) improves these models by aligning them with real-world observations, but marine biogeochemistry DA faces challenges due to model complexity, strong nonlinearity, and sparse, uncertain observations. Existing DA methods applied to marine biogeochemistry struggle to update unobserved variables effectively, while ensemble-based methods are computationally too expensive for high-complexity marine biogeochemistry models. This study demonstrates how machine learning (ML) can improve marine biogeochemistry DA by learning statistical relationships between observed and unobserved variables. We integrate ML-driven balancing schemes into a 1D prototype of a system used to forecast marine biogeochemistry in the North-West European Shelf seas. ML is applied to predict (i) state-dependent correlations from free-run ensembles and (ii), in an ``end-to-end'' fashion, analysis increments from an Ensemble Kalman Filter. Our results show that ML significantly enhances updates for previously not-updated variables when compared to univariate schemes akin to those used operationally. Furthermore, ML models exhibit moderate transferability to new locations, a crucial step toward scaling these methods to 3D operational systems. We conclude that ML offers a clear pathway to overcome current computational bottlenecks in marine biogeochemistry DA and that refining transferability, optimizing training data sampling, and evaluating scalability for large-scale marine forecasting, should be future research priorities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。