测试大模型对量子计算的理解极限,发现顶尖模型仍存认知短板。
Quantum-Audit: Evaluating the Reasoning Limits of LLMs on Quantum Computing
- 构建2700题量子计算测评集,涵盖专家出题、论文提取和含错设问
- 顶级模型最高达84%准确率,但面对专家题目平均低12分,安全类题仅73%
- 多数模型会接受错误前提,无法识别逻辑陷阱,适合评估AI推理能力
语言模型已成为量子计算教育与研究的实用工具,可用于总结技术论文、解释理论概念及回答领域最新进展问题。尽管现有基准已评估量子代码生成与电路设计能力,但对量子概念理解的系统性评测仍不足。Quantum-Audit通过2,700道涵盖核心量子计算主题的问题填补这一空白。我们评估了来自领先机构的26个模型。该基准包含1,000道专家撰写的题目、1,000道由大模型从研究论文中提取并经专家验证的题目,以及额外700道题目,包括350道开放问答题和350道含有错误前提的问题,以检验模型纠正错误假设的能力。人类参与者得分介于23%至86%之间,专家平均为74%。表现最佳的模型超越专家平均水平,Claude Opus 4.5 达到84%准确率,但顶级模型在专家题目上的平均准确率比自动生成题目低12个百分点。在高级主题上性能进一步下降,安全类问题准确率降至73%。此外,模型频繁接受并强化问题中嵌入的错误前提,此类关键推理任务的准确率低于66%。
原文摘要 · Abstract (English)
Language models have become practical tools for quantum computing education and research, from summarizing technical papers to explaining theoretical concepts and answering questions about recent developments in the field. While existing benchmarks evaluate quantum code generation and circuit design, their understanding of quantum computing concepts has not been systematically measured. Quantum-Audit addresses this gap with 2,700 questions covering core quantum computing topics. We evaluate 26 models from leading organizations. Our benchmark comprises 1,000 expert-written questions, 1,000 questions extracted from research papers using LLMs and validated by experts, plus an additional 700 questions including 350 open-ended questions and 350 questions with false premises to test whether models can correct erroneous assumptions. Human participants scored between 23% and 86%, with experts averaging 74%. Top-performing models exceeded the expert average, with Claude Opus 4.5 reaching 84% accuracy, though top models showed an average 12-point accuracy drop on expert-written questions compared to LLM-generated ones. Performance declined further on advanced topics, dropping to 73% on security questions. Additionally, models frequently accepted and reinforced false premises embedded in questions instead of identifying them, with accuracy below 66% on these critical reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。