arXiv:2507.21695physics.data-ancs.AI2025-07被引 3

构建物理领域大型评测基准,推动AI助力基础物理研究

Towards a Large Physics Benchmark

  • 三类题型结合概念、推导与开放问题,全面评估模型能力
  • 专家评分涵盖正确性、难度与意外性,确保评测科学性
  • 支持科研人员持续贡献新题,保持基准动态更新

我们提出一个由科学界共建共享的基准框架,用于评估、监控并引导大语言模型在基础物理领域的研发。基于科学理解与创造力的哲学理念,设计评分体系:每道题由专家评定正确性、难度与惊喜度。题目分为三类:(i)选择题考察概念理解,(ii)分析题需数学推导,(iii)开放任务要求复杂求解。当前数据集包含多样化示例,如高能物理事件分类任务(如四顶夸克信号)。为保证持续相关性,提出“活基准”机制,鼓励物理学家在发表论文时同步提交新问题。欢迎通过 http://www.physicsbenchmarks.org/ 贡献内容。期望该基准能推动有针对性的AI发展,切实助力基础物理研究。

原文摘要 · Abstract (English)

We introduce a benchmark framework developed by and for the scientific community to evaluate, monitor and steer large language model development in fundamental physics. Building on philosophical concepts of scientific understanding and creativity, we develop a scoring system in which each question is scored by an expert for its correctness, difficulty, and surprise. The questions are of three forms: (i) multiple-choice questions for conceptual understanding, (ii) analytical problems requiring mathematical derivation, and (iii) openended tasks requiring complex problem solving. Our current dataset contains diverse set of examples, including a machine learning challenge to classify high-energy physics events, such as the four top quark signal. To ensure continued relevance, we propose a living benchmark, where physicists contribute questions, for instance alongside new publications. We invite contributions via: http://www.physicsbenchmarks.org/. We hope that this benchmark will enable a targeted AI development that can make a meaningful contribution to fundamental physics research.

物理智能评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。