arXiv:2607.22067cs.CLcs.AI2026-07

用核反应堆操作员考试测试大模型,发现微调+检索能接近人类水平。

Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies

  • 用思维链和检索增强微调,提升模型在核能基础题上的表现。
  • 最优策略下通过8张试卷,正确率达80.23%,逼近人类及格线。
  • 不同训练阶段需用不同文本切分方式,不能直接复用。

在安全关键领域,语言模型的能力可信度需经行业标准检验。我们以美国核管会反应堆操作员通用基础考试(GFE)为基准,对一个310亿参数的开源多模态模型(Gemma 4 31B-IT)进行评估,按人类考生80%的及格标准逐卷评分,不作向上取整。评估集涵盖2015至2021年3月期全部考试,共14份试卷(7个压水堆PWR、7个沸水堆BWR),含697道题目。测试了八种配置:三种模型状态(原始模型、基于提炼思维链的监督微调SFT、检索增强微调RAFT),搭配三种检索条件(无检索、基于BM25在能源部基础手册中检索,采用固定大小与结构感知两种切块方式)。模型初始准确率为51.94%,未通过任何试卷。使用固定大小切块的SFT方案通过8/14份试卷,压水堆正确率达80.23%,总平均79.77%,置信区间跨越及格线。但切块粒度偏好随训练状态改变:未微调时结构感知更优,微调后则固定大小更佳,说明基于原始模型优化的切块策略无法直接用于微调后模型。RAFT整体比SFT低2.2至2.3个百分点,且在所有反应堆类型与切块组合中均落后。整个流程仅需单台工作站,运行时无需网络,结果接近操作员级工程基础掌握水平,但尚未稳定达到。

原文摘要 · Abstract (English)

Competence claims for a language model in a safety-critical domain are credible when measured against a standard the domain already enforces. We evaluate an open-weight 31-billion-parameter multimodal model (Gemma 4 31B-IT) on the U.S. Nuclear Regulatory Commission Reactor Operator Generic Fundamentals Examination (GFE), scoring it paper by paper against the 80% criterion applied to every human candidate, with no rounding up. The evaluation set is a census of every GFE administered at the March sitting from 2015 to 2021, giving seven pressurized water reactor (PWR) and seven boiling water reactor (BWR) papers and 697 scored items. Eight configurations cross three model states, the base model, supervised fine-tuning (SFT) on distilled chain-of-thought rationales and retrieval-augmented fine-tuning (RAFT), with three retrieval conditions, none and BM25 retrieval over the Department of Energy Fundamentals Handbooks under fixed-size and structure-aware chunking. Out of the box it answers 51.94% correctly and passes no paper. SFT with fixed-size chunking retrieval passes 8 of 14, reaching 80.23% on PWR items and 79.77% pooled, with a Wilson interval spanning the threshold. The preferred chunking granularity reverses with training state, structure-aware before fine-tuning and fixed-size after, so chunking optimized against a base model cannot be inherited by its fine-tuned descendant. RAFT trails SFT by 2.2 to 2.3 percentage points overall, and the deficit holds in all four reactor-type and chunking strata. The pipeline runs on one workstation with no network access at run time, and the result approaches operator-level command of engineering fundamentals without reliably achieving it.

大模型评测核能安全检索增强微调策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。