用国际奥赛题评估大模型科学推理能力,发现视觉理解是主要短板。
ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

- 构建涵盖13项奥赛的结构化评测集,专家审核确保题目质量。
- 顶尖模型在部分国际奥赛中获奖级评分,但化学与长程推理仍弱。
- 适合关注大模型科学思维、评测方法的科研人员参考。
前沿大模型面临评测饱和与数据污染问题,导致科学推理能力难以真实评估。本文提出 extsc{ScienceArena},一个基于物理、化学、生物领域十三项公开奥赛(包括 IPhO 2025–2026、IChO 2025–2026、IBO 2023、USAPhO 2026、USNCO 2025)的奥赛风格基准。其开放性、多步骤题目采用过程评分规则,难以人工准确打分。通过专家审核的数字化流程,将官方试题、图表、解答与评分标准转化为结构化数据,并由奥赛金牌获得者验证。为实现高效评估,以金牌选手评分作为基准,校准大模型作为评判者,验证表明两个强模型评分与专家总分误差不超过1分。分析显示,失败主要源于视觉理解、结构一致性与全局问题控制,而非术语缺失。对十四款近期大模型进行交叉求解评估,发现顶级模型在多项国际奥赛中达到奖牌水平评分,但化学领域及长周期逻辑一致性仍是核心瓶颈。提供在线演示:https://science-arena.onrender.com/
原文摘要 · Abstract (English)
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。