arXiv:2602.05857cs.AI2026-02

构建真实生物研究场景的推理测试,评估AI科学家的实验分析能力

BABE: Biology Arena BEnchmark

  • 基于真实论文构建,任务贴近科研实际
  • 要求模型进行因果推断与跨尺度推理
  • 适合评估生物AI系统科学思维水平

大型语言模型(LLMs)已从基础对话拓展至高级科学推理。然而,现有生物领域评测常无法衡量研究人员关键能力:将实验结果与背景知识结合以得出有意义结论的能力。为此,我们提出BABE(Biology Arena BEnchmark),一个全面的基准,用于评估生物AI系统的实验推理能力。BABE源自同行评审的研究论文和真实生物研究,确保任务反映实际科学探究的复杂性与交叉性。该基准挑战模型进行因果推理与跨尺度推断。它为评估AI系统是否能像实际科学家一样思考提供了可靠框架,从而更真实地衡量其在生物学研究中的潜力。

原文摘要 · Abstract (English)

The rapid evolution of large language models (LLMs) has expanded their capabilities from basic dialogue to advanced scientific reasoning. However, existing benchmarks in biology often fail to assess a critical skill required of researchers: the ability to integrate experimental results with contextual knowledge to derive meaningful conclusions. To address this gap, we introduce BABE(Biology Arena BEnchmark), a comprehensive benchmark designed to evaluate the experimental reasoning capabilities of biological AI systems. BABE is uniquely constructed from peer-reviewed research papers and real-world biological studies, ensuring that tasks reflect the complexity and interdisciplinary nature of actual scientific inquiry. BABE challenges models to perform causal reasoning and cross-scale inference. Our benchmark provides a robust framework for assessing how well AI systems can reason like practicing scientists, offering a more authentic measure of their potential to contribute to biological research.

生物AI实验推理科学智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。