arXiv:2502.12054cs.AI2025-02ACL被引 67

构建首个物理推理综合评测基准,揭示大模型在复杂物理问题上的短板。

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

  • 设计1200道题的基准集,含知识与推理题,分三难度层级。
  • 模型平均仅31.95%正确率,硬题表现远低于知识题。
  • 发现四大瓶颈:定理应用、过程理解、计算、条件分析。

大型语言模型在数学和逻辑推理方面表现出色,但现有评估忽视了需要物理定理与约束的物理推理任务。我们提出PhysReason,一个包含1200道题的综合性基准,其中25%为知识类问题,75%为推理类问题,且按难度分为易、中、难三级。题目平均需8.1步求解,难题达15.6步,体现其复杂性。我们设计了物理解答自动评分框架,支持答案级与步骤级评估。顶级模型如Deepseek-R1、Gemini-2.0-Flash-Thinking和o3-mini-high在答案级评估中表现不足60%,从知识题(75.11%)到难题(31.95%)显著下降。步骤级分析揭示四大关键瓶颈:物理定理应用、物理过程理解、计算能力与物理条件分析。PhysReason为评估大模型物理推理能力提供了全新、全面的基准。代码与数据将发布于https://dxzxy12138.github.io/PhysReason。

原文摘要 · Abstract (English)

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoning-based (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard requiring 15.6, reflecting the complexity of physics-based reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive step-level evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identified four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position PhysReason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models. Our code and data will be published at https:/dxzxy12138.github.io/PhysReason.

物理推理评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。