评测大模型在前沿物理研究中的自主探索能力,发现现有模型表现有限。
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

- 构建了涵盖5个物理子领域的全流程科研任务评测集
- 顶尖模型综合得分低于50,暴露与真实科研需求的显著差距
- 适合关注AI科学发现与自主研究的科研人员使用
智能体科学范式要求人工智能具备稳健推理与长周期自主探索能力。然而,当前科学评测仍局限于领域知识理解和复杂推理,无法评估真实科研中的探索性与流程复杂性。本文提出PRL-Bench(Physics Research by LLMs),一个面向理论与计算物理的研究导向评测基准。该基准基于2025年8月以来《物理评论快报》最新100篇论文构建,经领域专家验证,覆盖天体物理、凝聚态物理、高能物理、量子信息和统计物理五大理论与计算密集型子领域。每个任务均复现真实科研的核心特征:探索性问题设定、长周期工作流与可验证终点,重建了真实的科研推理过程与研究流程。对前沿模型的评估显示,性能依然受限,最佳综合得分低于50,揭示了当前大模型能力与真实科研需求之间的显著差距。PRL-Bench为评估下一代人工智能科学家提供了可靠测试平台,推动AI向自主科学发现迈进。
原文摘要 · Abstract (English)
The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain confined to domain knowledge comprehension and complex reasoning, failing to evaluate the exploratory nature and procedural complexity of real-world research. In this work, we present research-oriented evaluations in theoretical and computational physics, a natural testbed with comprehensive domain knowledge, complex reasoning, and verifiable end-to-end workflows without reliance on experiments. Here we introduce PRL-Bench (Physics Research by LLMs), a benchmark designed to systematically map the capability boundaries of LLMs in executing end-to-end physics research. Constructed from 100 curated papers from the latest issues of Physical Review Letters since August 2025 and validated by domain experts, PRL-Bench covers five major theory- and computation-intensive subfields of modern physics: astrophysics, condensed matter physics, high-energy physics, quantum information, and statistical physics. Each task in the benchmark is designed to replicate the core properties of authentic scientific research, including exploration-oriented formulation, long-horizon workflows, and objective verifiability, thereby reconstructing the essential reasoning processes and research workflows of real physics research. Evaluation across frontier models shows that performance remains limited, with the best overall score below 50, revealing a pronounced gap between current LLM capabilities and the demands of real scientific research. PRL-Bench serves a reliable testbed for accessing next generation AI scientists advancing AI systems toward autonomous scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。