首个面向量子领域的LLM评测基准,评估模型对复杂量子知识的理解能力。
QuantumBench: A Benchmark for Quantum Problem Solving
- 构建涵盖9个量子领域的800道多选题,覆盖真实科研材料。
- 首次系统评估主流LLM在量子任务中的表现,发现格式敏感性影响结果。
- 为量子研究中高效使用LLM提供可量化的评估工具,适合量子计算与AI交叉研究者。
大型语言模型已融入众多科学工作流程,加速数据分析、假设生成与设计空间探索。然而,通用基准难以反映特定领域对专业知识和符号表达的需求,这一差距在量子科学中尤为突出,因其包含非直观现象和高阶数学。本文提出QuantumBench,首个面向量子领域的语言模型评测基准,系统评估模型对量子知识的理解与应用能力。基于公开资料,我们整理约800道题目及其答案,覆盖九个量子科学方向,并构建为八选项多选题数据集。利用该基准,我们评估多个现有大模型在量子任务中的表现,分析其对问题格式变化的敏感性。QuantumBench旨在指导大模型在量子研究中的有效应用。
原文摘要 · Abstract (English)
Large language models are now integrated into many scientific workflows, accelerating data analysis, hypothesis generation, and design space exploration. In parallel with this growth, there is a growing need to carefully evaluate whether models accurately capture domain-specific knowledge and notation, since general-purpose benchmarks rarely reflect these requirements. This gap is especially clear in quantum science, which features non-intuitive phenomena and requires advanced mathematics. In this study, we introduce QuantumBench, a benchmark for the quantum domain that systematically examine how well LLMs understand and can be applied to this non-intuitive field. Using publicly available materials, we compiled approximately 800 questions with their answers spanning nine areas related to quantum science and organized them into an eight-option multiple-choice dataset. With this benchmark, we evaluate several existing LLMs and analyze their performance in the quantum domain, including sensitivity to changes in question format. QuantumBench is the first LLM evaluation dataset built for the quantum domain, and it is intended to guide the effective use of LLMs in quantum research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。