arXiv:2505.07889cs.CL2025-05被引 4

构建生物实验流程推理基准,提升大模型对实验逻辑的准确理解。

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science

  • 基于2.2万份真实实验协议构建大规模推理数据集
  • 10个主流大模型在定量与安全推理上表现显著下降
  • 提出ProAgent模型,显著提升实验流程理解能力

自主科学实验的实现受限于大模型对生物实验严格流程逻辑和精度要求的理解能力。为此,我们提出 extbf{BioProBench},一个面向生物实验流程推理的综合性资源。该基准建立在 extbf{BioProCorpus}之上,包含22,413份人工编写的实验协议。基于此语料库,我们系统构建了523,784个任务实例,既可作为大规模训练数据,也可作为具有新评估指标的严谨基准。对10个主流大模型的评估显示,尽管一般理解能力较高,但在需要深度推理、定量精确性和安全意识的任务上性能显著下降。为验证BioProCorpus的改进价值,我们开发了 extbf{ProAgent},基于该语料库,其显著提升了当前技术水平。项目代码与模型详见https://github.com/YuyangSunshine/bioprobench 和 https://huggingface.co/BioProBench。

原文摘要 · Abstract (English)

The realization of autonomous scientific experimentation is currently limited by LLMs' struggle to grasp the strict procedural logic and accuracy required by biological protocols. To address this fundamental challenge, we present \textbf{BioProBench}, a comprehensive resource for procedural reasoning in biology. BioProBench is grounded in \textbf{BioProCorpus}, a foundational collection of 22,413 human-written protocols. From this corpus, we systematically constructed a dataset of 523,784 task instances, offering both a large-scale training resource and a rigorous benchmark with novel metrics. Evaluating 10 mainstream LLMs, we find that while general comprehension is high, performance drops significantly on tasks demanding deep reasoning, quantitative precision, and safety awareness. To demonstrate the value of BioProCorpus in mitigating these issues, we developed \textbf{ProAgent}, grounded in our corpus, ProAgent substantially advances the state-of-the-art. https://github.com/YuyangSunshine/bioprobench and https://huggingface.co/BioProBench.

生物实验流程推理大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。