打造能自主研究前沿物理的AI科学家,突破长流程推理瓶颈。
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
- 用自适应多轨迹探索与分层记忆增强长程推理能力
- 在真实物理研究任务上得分51.08,显著优于现有模型
- 适合追求自主科研系统的研究人员和开发者
大型语言模型的推理与工具使用能力催生了代理科学,但前沿理论与计算物理研究仍具挑战,因需深厚领域知识、长周期推理及可靠数值计算。我们提出PRL-Bench,一个基于100篇《物理评论快报》论文构建的研究复现基准,将真实科研流程提炼为可追踪的任务,包含明确中间产物和多样化评估标准;每项任务被领域专家估算需超过六小时才能由专业博士生独立复现。评估显示现有代理在长流程任务中仍不可靠。为此,我们提出PhysMaster,一种结合自适应MCTS多轨迹探索与分层记忆的科学代理,以提升长周期鲁棒性与知识积累能力。PhysMaster在PRL-Bench上取得51.08的最高总分,优于Codex、OpenHands、OpenClaw和ReAct,在不同基线模型上相对提升达14.13%至93.38%。错误分析表明,其显著减少因长流程执行不完整导致的失败,但瓶颈仍在于物理知识与解析推理能力。PRL-Bench与PhysMaster共同为推进前沿物理自主AI研究提供了严格基准与有效系统。
原文摘要 · Abstract (English)
Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation. We introduce PRL-Bench, a research-reproduction benchmark adapted from 100 Physical Review Letters papers across major areas of modern physics. PRL-Bench distills realistic research workflows into traceable tasks with explicit intermediate artifacts and diverse evaluation rubrics; each task is estimated by domain experts to require more than six hours for a specialized PhD student to reproduce independently. Evaluations show that existing agents remain unreliable on extended research workflows. We therefore present PhysMaster, a scientific agent combining adaptive MCTS-based multi-trajectory exploration with hierarchical memory to improve long-horizon robustness and knowledge accumulation. PhysMaster achieves the highest overall PRL-Bench score of 51.08, outperforming Codex, OpenHands, OpenClaw, and ReAct, and yields relative improvements of 14.13 percent to 93.38 percent across backbone models. Error analysis shows that PhysMaster substantially reduces failures from incomplete long-horizon execution, while remaining bottlenecks lie in physics knowledge and analytical reasoning. Together, PRL-Bench and PhysMaster provide a rigorous benchmark and effective system for advancing autonomous AI research in frontier physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。