arXiv:2603.27646cs.CLhep-lat2026-03被引 5

构建物理论文复现基准,测试大模型端到端科研能力

PRBench: End-to-end Paper Reproduction in Physics Research

  • 设计30个真实物理论文任务,要求模型从理解到编码全程自主完成
  • 顶尖模型平均得分仅34%,全部模型零成功复现,代码与数据错误频发
  • 揭示公式误实现、仿真调试失败等系统性缺陷,适合评估科研自动化进展

由大语言模型驱动的AI代理展现出强大的推理与问题解决能力,可辅助科学任务如公式推导和代码生成。然而,这些代理能否可靠地完成从真实科学论文出发的端到端复现仍待验证。我们提出PRBench,一个包含30个专家精选任务的基准,覆盖物理学11个子领域。每个任务要求代理理解已发表论文的方法,从头实现相应算法,并生成与原始论文一致的定量结果。代理仅获任务说明与论文内容,在沙箱环境中运行。所有任务均由北京大学物理学院20多个研究组的领域专家提供,均基于真实发表论文,经端到端复现验证并配有真实结果与评分细则。通过代理化评估流程,我们在PRBench上评估多个编程代理,分析其在科学推理与执行关键维度的表现。表现最好的代理——OpenAI Codex(基于GPT-5.3-Codex)——平均得分为34%。所有代理端到端回调成功率为零,尤其在数据准确性和代码正确性上表现极差。我们进一步识别出系统性失败模式,包括公式实现错误、无法调试数值模拟以及虚构输出数据。总体而言,PRBench为评估向自主科研迈进的进展提供了严格基准。

原文摘要 · Abstract (English)

AI agents powered by large language models exhibit strong reasoning and problem-solving capabilities, enabling them to assist scientific research tasks such as formula derivation and code generation. However, whether these agents can reliably perform end-to-end reproduction from real scientific papers remains an open question. We introduce PRBench, a benchmark of 30 expert-curated tasks spanning 11 subfields of physics. Each task requires an agent to comprehend the methodology of a published paper, implement the corresponding algorithms from scratch, and produce quantitative results matching the original publication. Agents are provided only with the task instruction and paper content, and operate in a sandboxed execution environment. All tasks are contributed by domain experts from over 20 research groups at the School of Physics, Peking University, each grounded in a real published paper and validated through end-to-end reproduction with verified ground-truth results and detailed scoring rubrics. Using an agentified assessment pipeline, we evaluate a set of coding agents on PRBench and analyze their capabilities across key dimensions of scientific reasoning and execution. The best-performing agent, OpenAI Codex powered by GPT-5.3-Codex, achieves a mean overall score of 34%. All agents exhibit a zero end-to-end callback success rate, with particularly poor performance in data accuracy and code correctness. We further identify systematic failure modes, including errors in formula implementation, inability to debug numerical simulations, and fabrication of output data. Overall, PRBench provides a rigorous benchmark for evaluating progress toward autonomous scientific research.

科学自动化论文复现大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。