arXiv:2502.15815cs.LGastro-ph.CO2025-02被引 38

构建理论物理AI评测基准,测试大模型在高能与宇宙学问题上的推理能力。

Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics

  • 设计57道从本科到研究级的新题,覆盖高能物理与宇宙学。
  • 最新模型在难题上仍普遍无法求解,研究级题目平均正确率不足30%。
  • 适合关注AI辅助科研的物理学者与大模型评测研究人员。

我们提出一个用于评估AI解决理论物理问题能力的基准(TPBench),聚焦高能物理与宇宙学领域。该基准首版包含57道难度各异的问题,涵盖本科至研究级水平,且均为原创题目,未收录于公开题库。我们在多种开源与闭源语言模型(包括o3-mini、o1、DeepSeek-R1、GPT-4o及Llama、Qwen系列)上进行了评估。尽管近期模型性能显著提升,但研究级问题仍普遍无法解答。我们探讨了自动验证与评分的挑战,并分析了常见失败模式。当前顶尖模型对研究人员仍助益有限,但结果表明未来实现AI辅助理论物理研究具备可能性。本文还讨论主要障碍及应对策略。数据集、答案、模型表现与更新信息可在tpbench.org公开获取。

原文摘要 · Abstract (English)

We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology. The first iteration of our benchmark consists of 57 problems of varying difficulty, from undergraduate to research level. These problems are novel in the sense that they do not come from public problem collections. We evaluate our data set on various open and closed language models, including o3-mini, o1, DeepSeek-R1, GPT-4o and versions of Llama and Qwen. While we find impressive progress in model performance with the most recent models, our research-level difficulty problems are mostly unsolved. We address challenges of auto-verifiability and grading, and discuss common failure modes. While currently state-of-the art models are still of limited use for researchers, our results show that AI assisted theoretical physics research may become possible in the near future. We discuss the main obstacles towards this goal and possible strategies to overcome them. The public problems and solutions, results for various models, and updates to the data set and score distribution, are available on the website of the dataset tpbench.org.

理论物理AI评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。