为核工程中的机器学习模型提供不确定性量化方法,提升可信度。
Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?
- 区分物理模型与数据驱动模型的不确定性来源
- 揭示数据噪声、覆盖不足等导致预测误差的机制
- 适合关注机器学习可信性与安全性的核工程研究人员
机器学习在核工程各领域已广泛应用,其成功得益于深度学习突破、计算能力提升及开源工具普及。然而,这些经验成果常超越我们对算法本质的理解,尤其缺乏对不确定性量化(UQ)的重视。数据驱动的机器学习模型在预测时存在近似不确定性,源于数据噪声、数据覆盖不全、外推、模型结构缺陷及训练过程随机性等因素。本文旨在阐明机器学习不确定性量化的必要性,对比物理模型与数据驱动模型在不确定性概念上的差异,系统讨论两类模型的不确定性来源,并通过实例演示多种不确定性量化技术。最后,强调建立验证、确认与不确定性量化框架对提升机器学习可信度的重要性。
原文摘要 · Abstract (English)
Machine learning (ML) has been leveraged to tackle a diverse range of tasks in almost all branches of nuclear engineering. Many of the successes in ML applications can be attributed to the recent performance breakthroughs in deep learning, the growing availability of computational power, data, and easy-to-use ML libraries. However, these empirical successes have often outpaced our formal understanding of the ML algorithms. An important but under-rated area is uncertainty quantification (UQ) of ML. ML-based models are subject to approximation uncertainty when they are used to make predictions, due to sources including but not limited to, data noise, data coverage, extrapolation, imperfect model architecture and the stochastic training process. The goal of this paper is to clearly explain and illustrate the importance of UQ of ML. We will elucidate the differences in the basic concepts of UQ of physics-based models and data-driven ML models. Various sources of uncertainties in physical modeling and data-driven modeling will be discussed, demonstrated, and compared. We will also present and demonstrate a few techniques to quantify the ML prediction uncertainties. Finally, we will discuss the need for building a verification, validation and UQ framework to establish ML credibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。