评测大模型在核工程领域的知识与推理能力,发现其量化理解仍存短板。
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
- 构建混合流程生成1240组核工程问答对,覆盖事实、数值和语义三类任务。
- 大模型在事实题上表现较好,但数值推理和概念理解准确率显著偏低。
- 适合评估技术领域大模型能力,尤其关注科学推理的可靠性。
大语言模型(LLMs)在众多任务中表现出色,但在高度专业领域的可靠性仍面临挑战。核工程问题解决不仅需事实知识,还需定量推理和概念理解。为此,我们提出NuclearQAv2,一个用于评估大模型核工程知识的结构化基准。该基准包含约1,240个问答对,涵盖布尔型、数值型和语义型三类。通过结合专家编写、现有数据集及基于领域技术文献的LLM辅助生成,构建混合管道实现可扩展的基准构建。利用结构化提示进行自动问答生成与评分,支持高效评估。我们在NuclearQAv2上评估多种LLMs,发现模型在不同任务间表现差异明显:事实类问题表现良好,但定量推理与概念理解仍具挑战性。结果凸显多维度评估的重要性,并确立NuclearQAv2作为技术领域大模型能力评估的可扩展基准。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem solving often requires not only factual knowledge but also quantitative reasoning and conceptual understanding. To address the need for systematic evaluation in this domain, we introduce NuclearQAv2, a benchmark for assessing LLMs on nuclear engineering knowledge. The benchmark comprises approximately 1,240 question-answer pairs spanning three categories: boolean, numeric, and verbal. NuclearQAv2 is constructed using a hybrid pipeline that combines expert-authored questions, existing datasets, and LLM-assisted generation from domain-specific technical corpora. By leveraging structured prompting for both automated question generation and response evaluation, the proposed framework enables scalable benchmark construction and evaluation. We evaluate a diverse set of LLMs using NuclearQAv2 and observe substantial performance differences across task types. While the models generally perform well on factual questions, quantitative reasoning and conceptual understanding remain considerably more challenging. These results highlight the importance of multi-faceted evaluation frameworks and establish NuclearQAv2 as a scalable benchmark for assessing LLM capabilities in technical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。