241位专家共建398题高精度科学问答集,验证前沿模型真实水平。
Expert-validated STEM QA

- 由领域专家分四轮审核构建,确保问题与答案可验证且分布均衡
- 顶尖模型在该数据集上表现不足25%,体现真实评估价值
- 适合用于模型训练与科学推理能力评测,已开源部分数据
近期人工智能进步正推动数学、医学、材料科学等领域的突破。为促进这一进展,需高质量的评估数据集。现有STEM数据集存在诸多不足:模型性能趋于饱和、分类体系偏斜、多选题形式与科研实际脱节、答案错误率较高,主因是竞赛式采集和限时评审机制。本研究提出‘Expert-validated STEM QA’,一个由241位领域专家构建的高质量科学问答数据集(N=398),涵盖物理、化学、生物与数学。我们通过平衡分类体系、以质量为导向激励贡献者、多轮专家共识评审,并采用可验证的问答格式。实验表明,前沿模型在该数据集上的表现低于25%。在私有版本(N=2,000)上进行微调后,开源模型在HLE-verified数据集的STEM子集上性能相对基线提升15%(p=0.045),证明其训练潜力。部分数据已对社区开放。
原文摘要 · Abstract (English)
Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonomy distributions, (3) multiple choice question format that is misaligned with how scientists use AI in the real world, and (4) inaccurate answers and rationales partially led by a contest-based data collection and a time-bound review process. In this study, we present 'Expert-validated STEM QA', a high-quality, expert-validated STEM dataset (N=398) in Physics, Chemistry, Biology, and Mathematics, created by 241 domain experts. We (1) carefully designed a taxonomy with balanced distribution, (2) vetted question contributors with quality-driven incentive, (3) conducted multiple rounds of reviews with revisions validated by domain experts based on consensus, and (4) created the dataset in verifiable question and answer format. Our study demonstrated low performance ($<25\%$) of frontier AI models on the dataset as a benchmark. Post-training on a separate, private version of the dataset (N=2,000) increased performance of the open source model by $15\%$ relative to the baseline model (p=0.045) on the STEM subset of HLE-verified dataset, indicating potential utility of the dataset for model training. We have open-sourced a portion of our dataset for the AI research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。