构建动态科学评测集,可靠评估大模型的科学理解能力。
The Ever-Evolving Science Exam
- 构建超10万题的专家题库,覆盖5大学科500+子领域。
- 定期更新500题小样本集,避免数据泄露且降低评估成本。
- 可区分模型在科学与认知维度的优劣,适合模型评测者使用。
随着基础模型能力快速提升和广泛应用,评估其科学理解力变得愈发重要。现有科学评测基准虽在广度、覆盖范围和严谨性上取得进展,但仍面临两大挑战:数据泄露风险影响评测有效性,以及大规模测试导致的评估效率低下。为此,我们提出动态评测基准Ever-Evolving Science Exam(EESE),包含两个组件:1)一个非公开的EESE-Pool,包含超过10万道由专家构建的科学题目-答案对,覆盖5个学科和500多个子领域,通过多阶段流程确保广度、覆盖范围与严谨性;2)一个定期更新的500题子集EESE,经过采样与验证,实现抗泄露、低开销的评估。在32个开源与闭源模型上的实验表明,EESE能有效区分模型在科学领域和认知维度上的强弱。整体而言,EESE为科学评测提供了一种稳健、可扩展且面向未来的解决方案,真实衡量基础模型应对科学问题的能力。项目主页:https://github.com/aiben-ch/EESE。
原文摘要 · Abstract (English)
As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progress towards broad Range, wide Reach, and high Rigor, yet they often face two major challenges: data leakage risks that compromise benchmarking validity, and evaluation inefficiency due to large-scale testing. To address these issues, we introduce the Ever-Evolving Science Exam (EESE), a dynamic benchmark designed to reliably assess scientific capabilities in foundation models. Our approach consists of two components: 1) a non-public EESE-Pool with over 100K expertly constructed science instances (question-answer pairs) across 5 disciplines and 500+ subfields, built through a multi-stage pipeline ensuring Range, Reach, and Rigor, 2) a periodically updated 500-instance subset EESE, sampled and validated to enable leakage-resilient, low-overhead evaluations. Experiments on 32 open- and closed-source models demonstrate that EESE effectively differentiates the strengths and weaknesses of models in scientific fields and cognitive dimensions. Overall, EESE provides a robust, scalable, and forward-compatible solution for science benchmark design, offering a realistic measure of how well foundation models handle science questions. The project page is at: https://github.com/aiben-ch/EESE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。