构建专家级凝聚态理论基准,测试大模型物理推理能力。
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
- 由全球专家协作设计50道凝聚态理论难题,覆盖量子多体与统计力学。
- 顶尖模型仅解30%问题,平均正确率11.4%,多数难题涉及蒙特卡洛与张量网络。
- 发现模型常犯对称性错误,适合评估AI科研助手的物理理解力。
大语言模型在编程和数学求解上进展显著,但在硬科学前沿研究问题上的评估仍匮乏。为此,我们提出CMT-Benchmark,一个由全球专家团队共同设计的50个凝聚态理论问题数据集,涵盖量子多体与经典统计力学的分析与计算方法。问题涉及哈特里-福克、精确对角化、量子/变分蒙特卡洛、密度矩阵重整化群(DMRG)、量子/经典统计力学及模型构建。通过程序化方式验证答案,引入符号处理非对易算符的正则化机制进行机器评分。评估显示,前沿模型普遍表现不佳:最佳模型GPT5仅解决30%问题,17个模型平均准确率为11.4±2.1%。其中18题全模型未解,26题最多仅一模型解答。未解问题集中于量子蒙特卡洛、变分蒙特卡洛和DMRG,部分答案违反基本对称性或出现非物理解析结构。该基准可推动具备真正物理推理能力的AI研究助理发展。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce. To fill this gap, we present CMT-Benchmark, a dataset of 50 problems covering condensed matter theory (CMT) at the level of an expert researcher. Topics span analytical and computational approaches in quantum many-body, and classical statistical mechanics. The dataset was designed and verified by a panel of expert researchers from around the world. We built the dataset through a collaborative environment that challenges the panel to write and refine problems they would want a research assistant to solve, including Hartree-Fock, exact diagonalization, quantum/variational Monte Carlo, density matrix renormalization group (DMRG), quantum/classical statistical mechanics, and model building. We evaluate LLMs by programmatically checking solutions against expert-supplied ground truth. We developed machine-grading, including symbolic handling of non-commuting operators via normal ordering. They generalize across tasks too. Our evaluations show that frontier models struggle with all of the problems in the dataset, highlighting a gap in the physical reasoning skills of current LLMs. Notably, experts identified strategies for creating increasingly difficult problems by interacting with the LLMs and exploiting common failure modes. The best model, GPT5, solves 30\% of the problems; average across 17 models (GPT, Gemini, Claude, DeepSeek, Llama) is 11.4\pm2.1\%. Moreover, 18 problems are solved by none of the 17 models, and 26 by at most one. These unsolved problems span Quantum Monte Carlo, Variational Monte Carlo, and DMRG. Answers sometimes violate fundamental symmetries or have unphysical scaling dimensions. We believe this benchmark will guide development toward capable AI research assistants and tutors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。