构建实验室感知的化学语言模型评测基准,提升真实实验可靠性。
onepot-Bench 0: towards lab-aware in silico chemistry benchmarks
- 设计三类测试:化学计算、安全拒绝行为、反应结果预测。
- 使用自研实验数据评估催化剂选择与反应结果预测能力。
- 专为湿实验场景设计,避免训练数据泄露问题。
语言模型在实验室科学中日益重要,可完成实验规划、执行和事后分析。但精准评估其能力困难,因科学任务需结合问题解决与领域直觉。现有评估常忽视真实实验室中的可靠决策能力,且依赖可能已出现在模型训练数据中的公开数据。本文提出 onepot-Bench 0,一个专有的基准套件,用于评估语言模型在与湿实验相关合成化学能力方面的表现。该基准包含三个互补评估:ChemAbacus 测量无工具化学信息学素养与数值推理;SynthRefusal 分析在多种良性、受控及设计药物靶标下的安全与拒绝行为;SynthBench 使用我们实验室生成的私有实验数据,评估反应结果预测与催化剂选择能力。三者共同考察基础能力、可靠性与深层知识,均为实验室可靠运行所必需。
原文摘要 · Abstract (English)
Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution. onepot-Bench 0 comprises three complementary evaluations: ChemAbacus measures tool-free cheminformatics literacy and numerical reasoning; SynthRefusal characterizes safety and refusal behavior across a variety of benign, controlled, and designer-drug targets; and SynthBench evaluates reaction-outcome prediction and catalyst selection using private experimental data generated in our laboratory. Together, these evaluations probe basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。