arXiv:2509.26574cs.AIcond-mat.other2025-09被引 14

首个面向前沿物理研究的推理测试基准,揭示大模型在真实科研任务中的严重不足。

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

  • 构建71个由物理学家原创的真实科研难题,涵盖11个现代物理领域
  • 顶尖模型在完整任务上平均准确率仅5.7%,代码工具提升至约10%
  • 专为机器可验证、抗猜测设计,适合评估科学智能工具的实际能力

尽管大型语言模型在高中数学竞赛和编程任务中进展迅速,但它们能否有效应对前沿物理学研究中复杂的开放式挑战?更重要的是,物理学家希望大模型协助哪些推理任务?为回答这些问题,我们提出了CritPt(Complex Research using Integrated Thinking - Physics Test),首个针对未发表、研究级推理任务的基准,广泛覆盖凝聚态物理、量子物理、原子分子光学、天体物理、高能物理、数学物理、统计物理、核物理、非线性动力学、流体动力学及生物物理等现代物理领域。CritPt包含71个复合研究挑战,模拟入门级科研项目全貌,并分解为190个更简单的检查点任务以实现细粒度分析。所有问题均由50余位活跃物理研究员基于自身研究创作,经人工精校,确保答案不可猜测且机器可验证,采用高度定制化的自动化评分流水线评估复杂物理输出格式。结果显示,当前最先进的大模型虽在孤立检查点上展现初步潜力,但在完整研究级任务上仍表现欠佳:基础模型最高平均准确率为5.7%(GPT-5 高),配备代码工具后升至约10%。通过提供真实且标准化的评估,CritPt揭示了当前模型能力与真实科研需求间的巨大鸿沟,为发展基于科学的AI工具提供了坚实基础。

原文摘要 · Abstract (English)

While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason effectively through complex, open-ended challenges found in frontier physics research? And crucially, what kinds of reasoning tasks do physicists want LLMs to assist with? To address these questions, we present the CritPt (Complex Research using Integrated Thinking - Physics Test, pronounced "critical point"), the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics, astrophysics, high energy physics, mathematical physics, statistical physics, nuclear physics, nonlinear dynamics, fluid dynamics and biophysics. CritPt consists of 71 composite research challenges designed to simulate full-scale research projects at the entry level, which are also decomposed to 190 simpler checkpoint tasks for more fine-grained insights. All problems are newly created by 50+ active physics researchers based on their own research. Every problem is hand-curated to admit a guess-resistant and machine-verifiable answer and is evaluated by an automated grading pipeline heavily customized for advanced physics-specific output formats. We find that while current state-of-the-art LLMs show early promise on isolated checkpoints, they remain far from being able to reliably solve full research-scale challenges: the best average accuracy among base models is only 5.7%, achieved by GPT-5 (high), moderately rising to around 10% when equipped with coding tools. Through the realistic yet standardized evaluation offered by CritPt, we highlight a large disconnect between current model capabilities and realistic physics research demands, offering a foundation to guide the development of scientifically grounded AI tools.

AI推理物理研究基准测试科学智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。