arXiv:2608.25097cs.AIcs.MM2026-08

评测大模型解奥赛物理题能力,最强仅33.7%正确

PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

论文配图:PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?
图 1 · 摘自论文原文
  • 构建含1.16万道奥赛级物理题的多模态数据集
  • 最强模型答案准确率仅33.7%,远未达标
  • 提供中英双语步骤解析,适合评估推理链

评估(多模态)大语言模型在物理问题上的表现,需要反映专家级物理推理难度和广度的基准。现有物理基准存在两大局限:(1) 缺乏高难度数据集,(2) 未能全面覆盖视觉形式、知识点和逐步求解过程。因此,当前数据集上模型的表现可能无法真实反映其解决复杂物理问题的能力。为解决此问题,我们提出 PhysElite,一个大规模双语多模态奥赛级物理推理基准。PhysElite 包含 11,586 道奥赛级别问题,每道题均配有对应视觉图示、中英文双语逐步解答推导及最终答案。我们对 18 个开源与闭源多模态大模型进行了评测,发现即使最强模型答案准确率也仅达 33.7%。我们还进行步骤级过程评估,诊断模型在推理链中的失败环节。数据集已发布于 https://huggingface.co/datasets/physelite/PhysElite。

原文摘要 · Abstract (English)

Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present PhysElite, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. PhysElite contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese-English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. Our datasets are released at https://huggingface.co/datasets/physelite/PhysElite.

物理推理大模型评测多模态奥赛题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。