评测大模型在凝聚态物理中的解题能力,发现顶尖模型仅36分。
CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
- 聚焦凝聚态物理计算题,要求模型独立推导完整解答。
- 引入可扩展表达式编辑距离(SEED),实现细粒度评分。
- 结果显示最佳模型准确率仅28%,差距明显,适合科研评估。
我们提出CMPhysBench,一个用于评估大语言模型在凝聚态物理领域能力的新基准。该基准包含超过520道研究生级别、精心设计的题目,覆盖磁性、超导、强关联系统等代表性子领域及基础理论框架。为确保对解题过程的深入理解,所有题目均为计算类问题,要求模型独立生成完整求解过程。同时,我们采用基于树结构的表达式表示方法,引入可扩展表达式编辑距离(SEED)评分机制,实现非二元的细粒度部分得分,更精准衡量预测结果与标准答案之间的相似性。实验结果表明,即使是最先进的模型Grok-4,在该基准上平均SEED得分也仅为36,准确率仅28%,凸显出大模型在此实践性强、前沿性高的领域仍存在显著能力鸿沟。代码与数据集已公开于https://github.com/CMPhysBench/CMPhysBench。
原文摘要 · Abstract (English)
We introduce CMPhysBench, designed to assess the proficiency of Large Language Models (LLMs) in Condensed Matter Physics, as a novel Benchmark. CMPhysBench is composed of more than 520 graduate-level meticulously curated questions covering both representative subfields and foundational theoretical frameworks of condensed matter physics, such as magnetism, superconductivity, strongly correlated systems, etc. To ensure a deep understanding of the problem-solving process,we focus exclusively on calculation problems, requiring LLMs to independently generate comprehensive solutions. Meanwhile, leveraging tree-based representations of expressions, we introduce the Scalable Expression Edit Distance (SEED) score, which provides fine-grained (non-binary) partial credit and yields a more accurate assessment of similarity between prediction and ground-truth. Our results show that even the best models, Grok-4, reach only 36 average SEED score and 28% accuracy on CMPhysBench, underscoring a significant capability gap, especially for this practical and frontier domain relative to traditional physics. The code anddataset are publicly available at https://github.com/CMPhysBench/CMPhysBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。