arXiv:2508.00079cs.CLcs.AI2025-08中稿 · Findings of the As…被引 12

用多智能体验证提升大模型解物理题能力,效果显著。

PhysicsEval: Inference-Time Techniques to Improve the Reasoning Proficiency of Large Language Models on Physics Problems

  • 引入多智能体框架,让小模型逐轮验证大模型解题过程。
  • 在初始表现差的问题上,性能提升明显,最高达37.2%准确率增益。
  • 开源1.96万道物理题数据集,适合评测和改进推理模型。

物理学是人类智慧的基石,推动技术发展并深化对宇宙基本规律的理解。当前研究聚焦于解决物理问题这一关键自然语言推理任务。本文评估前沿大语言模型在数学与描述类物理问题上的表现,并应用多种推理时技术及代理框架提升模型性能。包括由小型语言模型代理以累积方式验证大模型提出的解法,并对各类技术效果进行对比分析。实验表明,当多智能体框架应用于模型初始表现较差的问题时,性能有显著提升。此外,本文提出一个新的物理问题评估基准 ${ m P{ iny HYSICS}E{ iny VAL}}$,包含从各类物理教科书获取的19,609道题目及其来自物理论坛和教育网站的正确解答。代码与数据已公开于 https://github.com/areebuzair/PhysicsEval。

原文摘要 · Abstract (English)

The discipline of physics stands as a cornerstone of human intellect, driving the evolution of technology and deepening our understanding of the fundamental principles of the cosmos. Contemporary literature includes some works centered on the task of solving physics problems - a crucial domain of natural language reasoning. In this paper, we evaluate the performance of frontier LLMs in solving physics problems, both mathematical and descriptive. We also employ a plethora of inference-time techniques and agentic frameworks to improve the performance of the models. This includes the verification of proposed solutions in a cumulative fashion by other, smaller LLM agents, and we perform a comparative analysis of the performance that the techniques entail. There are significant improvements when the multi-agent framework is applied to problems that the models initially perform poorly on. Furthermore, we introduce a new evaluation benchmark for physics problems, ${\rm P{\small HYSICS}E{\small VAL}}$, consisting of 19,609 problems sourced from various physics textbooks and their corresponding correct solutions scraped from physics forums and educational websites. Our code and data are publicly available at https://github.com/areebuzair/PhysicsEval.

大模型推理物理问题多智能体评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。