构建本科物理多模态推理基准,评估AI解题能力。
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
- 设计包含3304道题的多模态物理题库,涵盖8个子领域。
- 当前模型在该基准上最高仅51.6%准确率,尤其在多步推理中表现差。
- 适合关注物理推理、多模态模型评估的研究者使用。
物理问题求解是人工智能的挑战性领域,需整合概念理解、数学推理与物理图示解读能力。现有评测无法全面覆盖本科物理的广度与复杂性,而本科水平提供了严谨且标准化的评估基准。为此,我们提出PhysUniBench,一个大规模多模态基准,专门用于评估多模态大语言模型(MLLMs)在本科物理题上的推理能力。该基准包含3,304道涵盖8个主要物理子领域的题目,每题配有一张视觉图示。题目类型包括开放式与选择题,经过多轮迭代筛选与专家评审,并按五级难度系统分级。实验表明,当前模型在物理推理上面临巨大挑战,例如GPT-5在该基准上仅达51.6%准确率,尤其在多步推理和精确图示理解任务中表现不佳。PhysUniBench为推动科学领域人工智能发展提供了一个全面、严格的评估工具,旨在促进具备更强物理推理、问题解决与多模态理解能力的模型研发。
原文摘要 · Abstract (English)
Physics problem-solving is a challenging domain for AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Existing evaluations fail to capture the full breadth and complexity of undergraduate physics, whereas this level provides a rigorous yet standardized testbed for pedagogical assessment of multi-step physical reasoning. To this end, we present PhysUniBench, a large-scale multimodal benchmark designed to evaluate and improve the reasoning capabilities of multimodal large language models (MLLMs) specifically on undergraduate-level physics problems. PhysUniBench consists of 3,304 physics questions spanning 8 major sub-disciplines of physics, each accompanied by one visual diagram. The benchmark includes both open-ended and multiple-choice questions, systematically curated and difficulty-rated through an iterative process. The benchmark's construction involved a rigorous multi-stage process, including multiple roll-outs, expert-level evaluation, automated filtering of easily solved problems, and a nuanced difficulty grading system with five levels. Through extensive experiments, we observe that current models encounter substantial challenges in physics reasoning, where GPT-5 achieves only 51.6% accuracy in the PhysUniBench. These results highlight that current MLLMs struggle with advanced physics reasoning, especially on multi-step problems and those requiring precise diagram interpretation. By providing a broad and rigorous assessment tool, PhysUniBench aims to drive progress in AI for Science, encouraging the development of models with stronger physical reasoning, problem-solving skills, and multimodal understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。