arXiv:2601.21165cs.AIcs.CY2026-01被引 38

测试大模型解决顶尖科学难题的能力,涵盖奥赛与科研级问题。

FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

  • 设计两类任务:国际奥赛题与博士级科研子任务
  • 包含160道开源题目,覆盖量子电动力学至合成有机化学
  • 采用过程评分框架,评估科研解题全过程表现

我们提出FrontierScience,一个评估前沿语言模型在专家级科学推理方面能力的基准。现有科学基准多依赖选择题或已有知识,难以反映最新模型进展。FrontierScience通过两个互补赛道填补这一空白:(1) 奥赛赛道,包含国际奥赛级别题目(如IPhO、IChO、IBO);(2) 研究赛道,包含博士级开放性科研子任务。该基准共含数百道题目(其中160道已开源),覆盖物理、化学、生物多个子领域,从量子电动力学到合成有机化学。所有奥赛题由国际奥赛奖牌得主及国家队教练原创,确保难度、原创性与事实准确性。所有研究题由博士生、博士后或教授撰写并验证。针对研究任务,我们引入基于细粒度评分标准的评估框架,可全程评估模型解决科研任务的能力,而非仅评判最终答案。

原文摘要 · Abstract (English)

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-choice knowledge questions or already published information. FrontierScience addresses this gap through two complementary tracks: (1) Olympiad, consisting of international olympiad problems at the level of IPhO, IChO, and IBO, and (2) Research, consisting of PhD-level, open-ended problems representative of sub-tasks in scientific research. FrontierScience contains several hundred questions (including 160 in the open-sourced gold set) covering subfields across physics, chemistry, and biology, from quantum electrodynamics to synthetic organic chemistry. All Olympiad problems are originally produced by international Olympiad medalists and national team coaches to ensure standards of difficulty, originality, and factuality. All Research problems are research sub-tasks written and verified by PhD scientists (doctoral candidates, postdoctoral researchers, or professors). For Research, we introduce a granular rubric-based evaluation framework to assess model capabilities throughout the process of solving a research task, rather than judging only a standalone final answer.

科学推理大模型评测科研能力奥赛题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。