arXiv:2505.19099cs.AIphysics.ed-ph2025-05NeurIPS被引 29

测试大模型看图解物理题的能力,发现普遍不靠谱。

SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning

  • 构建75%题目必须靠读图才能答对的物理推理基准
  • 顶尖模型在该基准上准确率不足60%
  • 适合研究视觉-语言推理与认知偏见的学者

我们提出SeePhys,一个大规模多模态基准,用于评估大语言模型在物理问题上的推理能力,涵盖从中学到博士资格考级别的题目。基准覆盖物理学7个核心领域,包含21类高度异构的图表。与以往视觉仅起辅助作用不同,本基准中75%的问题为视觉关键型,需准确提取图像信息才能正确解答。大量实验表明,即使最先进的视觉推理模型(如Gemini-2.5-pro和o4-mini)在该基准上准确率也低于60%。结果揭示当前大语言模型在视觉理解方面存在根本性挑战:(i)图像解读与物理推理之间缺乏严格耦合;(ii)仍严重依赖文本线索作为认知捷径。

原文摘要 · Abstract (English)

We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In contrast to prior works where visual elements mainly serve auxiliary purposes, our benchmark features a substantial proportion of vision-essential problems (75%) that mandate visual information extraction for correct solutions. Through extensive evaluation, we observe that even the most advanced visual reasoning models (e.g., Gemini-2.5-pro and o4-mini) achieve sub-60% accuracy on our benchmark. These results reveal fundamental challenges in current large language models' visual understanding capabilities, particularly in: (i) establishing rigorous coupling between diagram interpretation and physics reasoning, and (ii) overcoming their persistent reliance on textual cues as cognitive shortcuts.

视觉推理物理问答大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。