评测15个大模型在量子力学任务上的表现,发现顶尖模型准确率超80%。
Evaluating Large Language Models on Quantum Mechanics: A Comparative Study Across Diverse Models and Tasks
- 对比15个模型在20类量子力学任务中的表现,分三档能力层级评估
- 旗舰模型平均准确率81%,数值计算仍是难点(仅42%正确)
- 工具增强可提效但效果不一,部分模型反而下降,需多轮验证
我们系统评估了大型语言模型在量子力学问题求解上的表现。研究涵盖来自五家厂商(OpenAI、Anthropic、Google、Alibaba、DeepSeek)的15个模型,分属三个能力层级,在20项任务上进行评估,包括推导、创造性问题、非标准概念和数值计算,共900个基准测试和75个工具增强测试。结果揭示清晰的层级分化:旗舰模型平均准确率达81%,分别领先中等模型(77%)和快速模型(67%)4个百分点和14个百分点。任务难度特征明显:推导类任务表现最好(平均92%,旗舰模型达100%),而数值计算仍最困难(仅42%)。工具增强在数值任务上效果各异:整体提升4.4个百分点,但耗时增加3倍,个别任务提升29个百分点,也有任务下降16个百分点。三次运行的可复现性分析显示平均方差6.3个百分点,旗舰模型(如GPT-5)表现极稳定(零方差),而专用模型需多轮评估。本工作贡献包括:(i)带自动验证的量子力学基准;(ii)层级性能差异的量化分析;(iii)工具增强权衡的实证研究;(iv)可复现性表征。所有任务、验证器与结果均公开发布。
原文摘要 · Abstract (English)
We present a systematic evaluation of large language models on quantum mechanics problem-solving. Our study evaluates 15 models from five providers (OpenAI, Anthropic, Google, Alibaba, DeepSeek) spanning three capability tiers on 20 tasks covering derivations, creative problems, non-standard concepts, and numerical computation, comprising 900 baseline and 75 tool-augmented assessments. Results reveal clear tier stratification: flagship models achieve 81\% average accuracy, outperforming mid-tier (77\%) and fast models (67\%) by 4pp and 14pp respectively. Task difficulty patterns emerge distinctly: derivations show highest performance (92\% average, 100\% for flagship models), while numerical computation remains most challenging (42\%). Tool augmentation on numerical tasks yields task-dependent effects: modest overall improvement (+4.4pp) at 3x token cost masks dramatic heterogeneity ranging from +29pp gains to -16pp degradation. Reproducibility analysis across three runs quantifies 6.3pp average variance, with flagship models demonstrating exceptional stability (GPT-5 achieves zero variance) while specialized models require multi-run evaluation. This work contributes: (i) a benchmark for quantum mechanics with automatic verification, (ii) systematic evaluation quantifying tier-based performance hierarchies, (iii) empirical analysis of tool augmentation trade-offs, and (iv) reproducibility characterization. All tasks, verifiers, and results are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。