arXiv:2505.15929cs.AI2025-05被引 32

首个大规模物理推理测试,揭示大模型在真实世界物理理解上的严重短板。

PhyX: Does Your Model Have the "Wits" for Physical Reasoning?

  • 构建6大物理领域3000题的多模态推理集,覆盖25个子领域
  • 顶尖模型最高仅45.8%准确率,远低于人类专家
  • 适合评估视觉-物理联合推理能力的研究者使用

现有基准无法捕捉智能的核心能力:物理推理,即整合领域知识、符号推理与现实约束理解的能力。为此,我们提出PhyX:首个面向视觉场景中物理基础推理的大规模基准。PhyX包含3000道精心设计的多模态问题,覆盖6类推理类型、25个子领域及6大核心物理领域:热力学、电磁学、力学、现代物理、光学与波与声学。全面评估显示,即使最先进的模型在物理推理上仍表现不佳:GPT-4o、Claude3.7-Sonnet和GPT-o4-mini分别仅达到32.5%、42.2%和45.8%准确率,与人类专家差距超过29%。分析发现当前模型存在三大缺陷:过度依赖记忆化学科知识、过度依赖数学表达形式、以及仅基于表面视觉模式匹配而非真正的物理理解。我们通过细粒度统计、案例研究和多重评估范式深入剖析物理推理能力。为保证可复现性,我们基于VLMEvalKit等通用工具包实现兼容评估协议,支持一键评估。更多详情见项目页:https://phyx-bench.github.io/

原文摘要 · Abstract (English)

Existing benchmarks fail to capture a crucial aspect of intelligence: physical reasoning, the integrated ability to combine domain knowledge, symbolic reasoning, and understanding of real-world constraints. To address this gap, we introduce PhyX: the first large-scale benchmark designed to assess models capacity for physics-grounded reasoning in visual scenarios. PhyX includes 3K meticulously curated multimodal questions spanning 6 reasoning types across 25 sub-domains and 6 core physics domains: thermodynamics, electromagnetism, mechanics, modern physics, optics, and wave\&acoustics. In our comprehensive evaluation, even state-of-the-art models struggle significantly with physical reasoning. GPT-4o, Claude3.7-Sonnet, and GPT-o4-mini achieve only 32.5%, 42.2%, and 45.8% accuracy respectively-performance gaps exceeding 29% compared to human experts. Our analysis exposes critical limitations in current models: over-reliance on memorized disciplinary knowledge, excessive dependence on mathematical formulations, and surface-level visual pattern matching rather than genuine physical understanding. We provide in-depth analysis through fine-grained statistics, detailed case studies, and multiple evaluation paradigms to thoroughly examine physical reasoning capabilities. To ensure reproducibility, we implement a compatible evaluation protocol based on widely-used toolkits such as VLMEvalKit, enabling one-click evaluation. More details are available on our project page: https://phyx-bench.github.io/.

物理推理多模态评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。