升级版生物科研AI评测基准,更贴近真实研究场景。
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research

- 构建近1900个真实科研任务,评估AI执行实际研究的能力。
- 新基准难度显著提升,模型准确率下降26%至46%。
- 适合开发科学智能工具的研究者使用,推动AI科研落地。
AI加速科学发现的前景持续乐观。当前应用涵盖科学数据训练的基础模型、自主生成假设的智能体系统,以及人工智能驱动的自动化实验室。衡量AI在科学领域进展的需求必须同步加强,并逐步转向更真实的科研能力评估,而不仅限于记忆或推理。此前提出的语言代理生物学基准(LAB-Bench)是首次尝试。本文推出其演进版本LABBench2,用于评估AI在真实科研中完成有用任务的能力。LABBench2包含近1900个任务,延续前版框架,但测试环境更贴近现实。我们评估了前沿模型表现,发现尽管整体能力有所提升,但LABBench2带来的难度跃升显著——各子任务中模型准确率下降26%至46%。该基准延续LAB-Bench作为科学智能能力评估的行业标准地位,助力推动科研类AI工具发展。为促进社区使用与开发,任务数据集已公开于https://huggingface.co/datasets/futurehouse/labbench2,评估工具链开源于https://github.com/EdisonScientific/labbench2。
原文摘要 · Abstract (English)
Optimism for accelerating scientific discovery with AI continues to grow. Current applications of AI in scientific research range from training dedicated foundation models on scientific data to agentic autonomous hypothesis generation systems to AI-driven autonomous labs. The need to measure progress of AI systems in scientific domains correspondingly must not only accelerate, but increasingly shift focus to more real-world capabilities. Beyond rote knowledge and even just reasoning to actually measuring the ability to perform meaningful work. Prior work introduced the Language Agent Biology Benchmark LAB-Bench as an initial attempt at measuring these abilities. Here we introduce an evolution of that benchmark, LABBench2, for measuring real-world capabilities of AI systems performing useful scientific tasks. LABBench2 comprises nearly 1,900 tasks and is, for the most part, a continuation of LAB-Bench, measuring similar capabilities but in more realistic contexts. We evaluate performance of current frontier models, and show that while abilities measured by LAB-Bench and LABBench2 have improved substantially, LABBench2 provides a meaningful jump in difficulty (model-specific accuracy differences range from -26% to -46% across subtasks) and underscores continued room for performance improvement. LABBench2 continues the legacy of LAB-Bench as a de facto benchmark for AI scientific research capabilities and we hope that it continues to help advance development of AI tools for these core research functions. To facilitate community use and development, we provide the task dataset at https://huggingface.co/datasets/futurehouse/labbench2 and a public eval harness at https://github.com/EdisonScientific/labbench2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。