arXiv:2604.09836cs.AIcs.CL2026-04被引 1

新基准测试评估AI在理化生数领域的科学推理能力,发现当前模型仍有巨大提升空间。

COMPOSITE-Stem

论文配图:COMPOSITE-Stem
图 1 · 摘自论文原文
  • 设计70个专家命题,融合精确匹配与评审标准,用大模型评判输出质量。
  • 顶尖模型仅达21%正确率,说明现有AI在复杂科学任务上仍不成熟。
  • 适合研究科学智能、智能体评估及跨学科应用的学者参考。

AI智能体在加速科学发现方面潜力巨大,但缺乏前沿评测阻碍了其在真实工作流中的应用。尽管专家撰写的基准已有效衡量AI推理能力,但多数已趋于饱和,仅能评估受限输出。为此,我们推出COMPOSITE-STEM,一个由博士级研究人员精心设计的70个理、化、生、数领域任务基准。该基准结合精确匹配评分与标准评审体系,并采用大模型作为评审员的评分机制,实现对科学有意义输出的灵活评估。我们使用改进的多模态Terminus-2智能体,在Harbor智能体评估框架下评测四种前沿模型。表现最佳模型仅达21%正确率,表明COMPOSITE-STEM能捕捉当前智能体尚未达到的能力。所有任务均在授权下开源,以支持可复现性并推动科学智能研究。

原文摘要 · Abstract (English)

AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have proven effective at measuring AI reasoning, but most at this stage have become saturated and only measure performance on constrained outputs. To help address this gap, we introduce COMPOSITE-STEM, a benchmark of 70 expert-written tasks in physics, biology, chemistry, and mathematics, curated by doctoral-level researchers. Our benchmark combines exact-match grading and criterion-based rubrics with an LLM-as-a-jury grading protocol, allowing more flexible assessment of scientifically meaningful outputs. Using an adapted multimodal Terminus-2 agent harness within the Harbor agentic evaluation framework, we evaluate four frontier models. The top-performing model achieves 21%, demonstrating that COMPOSITE-STEM captures capabilities beyond current agent reach. All tasks are open-sourced with contributor permission to support reproducibility and to promote additional research towards AI's acceleration of scientific progress in these domains.

科学智能智能体评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。