用VAE生成数据提升核工程预测精度并降低不确定性
An Investigation on Machine Learning Predictive Accuracy Improvement and Uncertainty Reduction using VAE-based Data Augmentation
- 用变分自编码器生成合成数据扩充训练集
- 模型预测误差降低,置信区间更紧凑,不确定性减少
- 适合数据稀缺领域的机器学习建模者参考
超快计算与大容量内存、机器学习算法进步及大规模数据集的出现,使多个工程领域迎来突破性进展。然而,核工程面临数据稀缺挑战,因核系统实验成本高、耗时长。深度生成学习可通过学习真实数据分布生成合成样本,显著扩充数据集以训练更精准的预测模型。本研究评估基于变分自编码器(VAE)的生成式数据增强对深度神经网络(DNN)预测准确性的提升效果,并采用贝叶斯神经网络(BNN)和约等预测(CP)量化预测不确定性。实验基于NUPEC沸水堆全尺寸细网格束测试(BFBT)基准的TRACE稳态空泡份额仿真数据。结果表明,使用VAE增强训练数据后,DNN模型预测精度提高,预测置信区间更紧致,预测不确定性显著降低。
原文摘要 · Abstract (English)
The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. We found that augmenting the training dataset using VAEs has improved the DNN model's predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。