通过神经元分析,让大模型在核反应堆安全中的决策过程可解释。
Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications
- 用低秩适配微调模型,识别出核领域专属的稀疏神经元
- 集体关闭这些神经元会显著降低任务表现,证明其关键作用
- 适合关注高安全要求AI可信性的核工程与AI交叉研究者
将大语言模型(LLMs)应用于核工程等高安全领域,需深入理解其内部推理机制。本文以沸水反应堆系统为案例,采用低秩适配(LoRA)技术将通用模型Gemma-3-1b-it微调至核领域。通过对比基础模型与微调后模型的神经元激活模式,发现少数神经元在适应过程中行为发生显著变化。利用神经元屏蔽技术探究其因果作用:单独屏蔽多数神经元未产生统计显著影响,但集体关闭这些神经元组导致任务性能显著下降。定性分析表明,屏蔽后模型生成技术细节的能力减弱。本研究提出一种可追踪领域知识到可验证神经回路的解释方法,为实现核级人工智能保障提供路径,应对核监管框架(如10 CFR 50 Appendix B)对验证与确认的要求,推动安全关键场景中AI的应用。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) into safety-critical domains, such as nuclear engineering, necessitates a deep understanding of their internal reasoning processes. This paper presents a novel methodology for interpreting how an LLM encodes and utilizes domain-specific knowledge, using a Boiling Water Reactor system as a case study. We adapted a general-purpose LLM (Gemma-3-1b-it) to the nuclear domain using a parameter-efficient fine-tuning technique known as Low-Rank Adaptation. By comparing the neuron activation patterns of the base model to those of the fine-tuned model, we identified a sparse set of neurons whose behavior was significantly altered during the adaptation process. To probe the causal role of these specialized neurons, we employed a neuron silencing technique. Our results demonstrate that while silencing most of these specialized neurons individually did not produce a statistically significant effect, deactivating the entire group collectively led to a statistically significant degradation in task performance. Qualitative analysis further revealed that silencing these neurons impaired the model's ability to generate detailed, contextually accurate technical information. This paper provides a concrete methodology for enhancing the transparency of an opaque black-box model, allowing domain expertise to be traced to verifiable neural circuits. This offers a pathway towards achieving nuclear-grade artificial intelligence (AI) assurance, addressing the verification and validation challenges mandated by nuclear regulatory frameworks (e.g., 10 CFR 50 Appendix B), which have limited AI deployment in safety-critical nuclear operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。