arXiv:2410.00929cs.AIcs.LG2024-10被引 6

用知识引导的LLM框架,高效分类核电站停机起因事件

A Knowledge-Informed Large Language Model Framework for U.S. Nuclear Power Plant Shutdown Initiating Event Classification for Probabilistic Risk Assessment

  • 先用44个关键词模式筛掉97%非停机事件,再用BERT模型分类
  • 在10,928条事件数据上,分类准确率达93.4%
  • 适合核能风险评估、安全分析等领域的研究人员

识别和分类停机起因事件(SDIEs)对开展核电站低功率停机概率风险评估至关重要。现有计算方法因缺乏大规模标注数据、事件类型不平衡及标签噪声等问题,难以取得理想效果。为此,我们提出一种混合流程:第一阶段利用44个基于六类SDIE的关键词与短语构建文本模式,通过向量化生成高可分特征,用简单二分类器预筛非SDIE;第二阶段构建基于BERT的大语言模型(LLM),在大规模语料自监督预训练后,微调于SDIE数据集,实现四类事件分类。在包含10,928个事件的数据集上评估,结果表明预筛阶段可排除超过97%非SDIE,LLM平均准确率达93.4%。

原文摘要 · Abstract (English)

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing low power shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning mode to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F1 score, and average accuracy. The results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 93.4% for SDIE classification.

核能安全LLM应用事件分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。