arXiv:2509.07894cs.AI2025-09被引 19

首个面向高中生物理奥赛的开源评测基准,可直接对比大模型与人类表现。

HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?

  • 构建2024-2025年13套奥赛真题,覆盖图文混合题型
  • 按官方评分标准逐步打分,实现与真人考官一致的评估
  • 首次用奖牌等级量化模型性能,适合评估物理推理能力

近年来,(M)LLMs在物理能力方面的表现引发关注。然而,现有物理评测基准存在两大缺陷:既未系统覆盖真实物理奥赛(如国际和区域级竞赛),也无法与人类表现直接比较。为此,我们提出HiPhO——首个专为高中生物理奥赛设计、具备人类对齐评估的基准。其三大创新包括:(1)全面数据:整合2024-2025年13套最新奥赛试卷,涵盖国际与区域赛事,题型从纯文本到图示题;(2)专业评估:采用官方评分标准,在答案与解题步骤层面进行细粒度评分,确保评估质量与领域专业性;(3)人类对标:依据官方奖牌线为模型分配金、银、铜牌,实现与人类参赛者直接对比。大规模评估30个先进(M)LLMs发现:开源多模态模型普遍仅达或低于铜牌水平;开源语言模型已展现潜力,部分获金牌;闭源推理型多模态模型可达6至12枚金牌;但多数模型仍距满分有显著差距。结果揭示了开源模型与顶尖学生间的性能鸿沟,凸显闭源模型的强大推理能力及未来改进空间。HiPhO现已开源,代码与公开排行榜见https://github.com/SciYu/HiPhO 及 https://phyarena.github.io/。

原文摘要 · Abstract (English)

Recently, the physical capabilities of (M)LLMs have garnered increasing attention. However, existing benchmarks for physics suffer from two major gaps: they neither provide systematic and up-to-date coverage of real-world physics competitions such as physics Olympiads, nor enable direct performance comparison with humans. To bridge these gaps, we present HiPhO, the first benchmark dedicated to high school physics Olympiads with human-aligned evaluation. Specifically, HiPhO highlights three key innovations. (1) Comprehensive Data: It compiles 13 latest Olympiad exams from 2024-2025, spanning both international and regional competitions, and covering mixed modalities that encompass problems spanning text-only to diagram-based. (2) Professional Evaluation: We adopt official marking schemes to perform fine-grained grading at both the answer and step level, fully aligned with human examiners to ensure high-quality and domain-specific evaluation. (3) Comparison with Human Contestants: We assign gold, silver, and bronze medals to models based on official medal thresholds, thereby enabling direct comparison between (M)LLMs and human contestants. Our large-scale evaluation of 30 state-of-the-art (M)LLMs shows that: across 13 exams, open-source MLLMs mostly remain at or below the bronze level; open-source LLMs show promising progress with multiple golds; closed-source reasoning MLLMs can achieve 6 to 12 gold medals; and most models still have a significant gap from full marks. These results highlight the performance gap between open-source models and top students, the strong reasoning abilities of closed-source models, and the remaining room for improvement. HiPhO, a human-aligned Olympiad benchmark for multimodal physical reasoning, is open-source at https://github.com/SciYu/HiPhO with a public leaderboard at https://phyarena.github.io/.

物理推理大模型评测奥赛基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。