评测大模型在真实实验流程修改中的推理能力,提升科研工具实用性。
BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

- 基于科学家真实修改的实验流程构建任务,确保评估贴近实际
- 最佳模型得分59.2%,多数模型在34%至47%之间,仍有提升空间
- 适合关注生物实验自动化与模型辅助科研的研究者
我们提出BenchBench-Protocol,一个包含149个协议修改任务的基准测试,数据源自科学家在真实实验中对已发表流程的实际修改。适应已有协议开展新实验是湿实验科学家的常规任务,正确修改需考虑前期决策和后续步骤。现有生命科学基准多采用专家设计的开放性题目,而本基准通过对比原始协议与科学家修改版本生成任务,提供查询依据和加权评分标准。数据来自9个湿实验生物学领域的96个原始协议,所有任务经领域专家评审后保留高质项。评估了九个闭源与开源模型,Claude Opus 5表现最优,得分为59.2%标准化评分,其余模型在34.1%至47.1%之间,且最优模型在十次尝试中仍未饱和。随着大模型在生命科学研究中日益重要,对其在常规湿实验任务上的表现进行评估愈发关键。BenchBench-Protocol不仅为湿实验推理提供可靠评估,也证明了真实实验经验对构建基准任务的价值。
原文摘要 · Abstract (English)
We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recovered from modifications that scientists made to published protocols during real experimental work. Adapting a published protocol to a new experiment is a routine task for a wet-lab scientist, and a correct modification requires accounting for prior choices and downstream steps. Recent life-science benchmarks have moved toward open-ended, rubric-graded tasks, but tasks are typically elicited from experts rather than reconstructed from real-world modifications. BenchBench-Protocol tasks are derived from differences between a published protocol and a version a scientist modified, which provides the basis for the query and the weighted rubric elements for a correct response. The benchmark draws from 96 source protocols across nine domains of wet-lab biology and only includes tasks rated highly after review by domain experts. We evaluate nine closed and open models; Claude Opus 5 scores highest at 59.2% normalized rubric score, with other models between 34.1% and 47.1%, and the benchmark remains unsaturated when taking the best of ten attempts. As models are increasingly helpful in life-sciences research, evaluating them on routine wet-lab tasks becomes correspondingly important. We present BenchBench-Protocol as both a grounded assessment of wet-lab reasoning and evidence for the utility of real-world experiments to construct benchmark tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。