用插值提升数据精度,三种模型对比优化核设施烟羽辐射剂量估算。
Interpolation-Driven Machine Learning Approaches for Plume Shine Dose Estimation: A Comparison of XGBoost, Random Forest, and TabNet
- 通过保形插值构建高分辨率训练数据,增强模型输入质量。
- XGBoost在所有场景中表现最佳,误差低于其他模型15%以上。
- 树模型专注关键几何参数,而TabNet更均衡利用多变量特征。
尽管机器学习在代理建模中取得成功,但在辐射剂量评估中受限于安全约束、数据稀缺及物理主导系统架构选择困难。本文以烟羽闪烁剂量快速估算为实例,该任务对核电设施安全与应急响应至关重要,传统光子输运计算成本过高。研究基于pyDOSEIA工具包生成17种伽马发射核素在不同下风距离、释放高度和大气稳定度条件下的离散剂量数据集,采用保形插值方法构建高分辨率训练数据。比较了两种树模型(随机森林、XGBoost)和一种深度学习模型(TabNet)的预测性能与数据分辨率敏感性。所有模型在插值后数据上均优于原始离散数据;其中XGBoost始终表现最优。基于排列重要性(树模型)和注意力特征归因(TabNet)的可解释性分析表明,性能差异源于模型对输入特征的使用方式:树模型主要依赖释放高度、稳定度和下风距离等主导几何-扩散特征,将核素种类视为次要输入;而TabNet在多个变量间分配注意力更均匀。为支持实际部署,开发了基于Web的图形界面,实现情景交互评估,并可与光子输运参考结果透明比对。
原文摘要 · Abstract (English)
Despite the success of machine learning (ML) in surrogate modeling, its use in radiation dose assessment is limited by safety-critical constraints, scarce training-ready data, and challenges in selecting suitable architectures for physics-dominated systems. Within this context, rapid and accurate plume shine dose estimation serves as a practical test case, as it is critical for nuclear facility safety assessment and radiological emergency response, while conventional photon-transport-based calculations remain computationally expensive. In this work, an interpolation-assisted ML framework was developed using discrete dose datasets generated with the pyDOSEIA suite for 17 gamma-emitting radionuclides across varying downwind distances, release heights, and atmospheric stability categories. The datasets were augmented using shape-preserving interpolation to construct dense, high-resolution training data. Two tree-based ML models (Random Forest and XGBoost) and one deep learning (DL) model (TabNet) were evaluated to examine predictive performance and sensitivity to dataset resolution. All models showed higher prediction accuracy with the interpolated high-resolution dataset than with the discrete data; however, XGBoost consistently achieved the highest accuracy. Interpretability analysis using permutation importance (tree-based models) and attention-based feature attribution (TabNet) revealed that performance differences stem from how the models utilize input features. Tree-based models focus mainly on dominant geometry-dispersion features (release height, stability category, and downwind distance), treating radionuclide identity as a secondary input, whereas TabNet distributes attention more broadly across multiple variables. For practical deployment, a web-based GUI was developed for interactive scenario evaluation and transparent comparison with photon-transport reference calculations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。