将生物地球化学模型嵌入神经网络,提升土壤碳预测精度与机制理解。
Biogeochemistry-Informed Neural Network (BINN) for Improving Accuracy of Model Prediction and Scientific Understanding of Soil Organic Carbon
- 融合过程模型与神经网络,用数据驱动方式学习土壤碳循环机制。
- 在25,925个观测点上预测6大碳循环过程,相关性达0.86,精度接近贝叶斯方法。
- 计算效率比传统方法提升50倍以上,适合大规模土壤碳研究。
随着大规模观测数据和人工智能技术的发展,提升对全球碳循环等生物地球化学过程的理解成为可能。然而,从大数据中提取机制性知识仍具挑战。本文提出一种生物地球化学信息神经网络(BINN),将向量化的过程型土壤碳循环模型(即社区土地模型5版,CLM5)无缝集成到神经网络结构中,以从大数据中解析土壤有机碳(SOC)储存的调控机制。在参数恢复实验中,BINN在合成数据上表现出高精度。通过引入蒙特卡洛丢弃生成后验分布,证明其可有效量化参数不确定性。利用BINN对美国本土25,925个土壤有机碳剖面进行六项主要过程预测,并与基于贝叶斯推断的PRODA方法结果对比,空间格局一致性良好(平均相关系数=0.86),表明其机制捕捉能力与经典贝叶斯方法相当。此外,该框架使计算效率较PRODA提升超50倍。结论表明,BINN是一种高效整合人工智能、大数据与过程模型的土壤碳循环研究新范式。
原文摘要 · Abstract (English)
The increasing availability of large-scale observational data and the rapid development of artificial intelligence (AI) provide unprecedented opportunities to enhance our understanding of the global carbon cycle and other biogeochemical processes. However, retrieving mechanistic knowledge from these large-scale data remains a challenge. Here, we develop a Biogeochemistry-Informed Neural Network (BINN) that seamlessly integrates a vectorized process-based soil carbon cycle model (i.e., Community Land Model version 5, CLM5) into a neural network (NN) structure to examine mechanisms governing soil organic carbon (SOC) storage from big data. BINN demonstrates high accuracy in retrieving biogeochemical parameter values from synthetic data in a parameter recovery experiment. Furthermore, by incorporating Monte Carlo (MC) dropout to generate posterior distributions, we demonstrate that BINN can effectively quantify uncertainty in estimated parameters. We use BINN to predict six major processes (or components in process-based models) regulating the soil carbon cycle from 25,925 observed SOC profiles across the contiguous US and compare them with the same processes previously retrieved by a Bayesian inference-based PROcess-guided deep learning and DAta-driven modeling (PRODA) approach. The good agreement between the spatial patterns retrieved by BINN and PRODA (average correlation coefficient = 0.86) suggests that BINN's ability of capturing mechanistic knowledge is consistent with the established Bayesian-based methods. Additionally, the integration of neural networks and process-based models in BINN improves computational efficiency by more than 50 times over PRODA. We conclude that BINN is an efficient framework that harnesses the power of both AI, large-scale data, and process-based modeling to understand large scale soil carbon cycle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。