测试AI在完整生物研究任务中的表现,发现其处理大数据和多步骤分析能力有限。
BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks

- 设计模拟科学家委托的生物研究任务,评估AI从原始数据到结果的全流程执行能力。
- 20个任务中最高得分仅0.48,大样本和多步骤任务表现显著下降。
- 顶尖AI虽效率高、成本低,但整体仍难胜任复杂科研任务,适合研究人员参考。
人工智能有望通过自动化计算分析加速生物学研究,但目前尚无系统性评估其在完整研究尺度下的表现。本文提出BixBench3基准,衡量AI代理从原始生物数据到科学结论的全过程处理能力。任务设计模拟科学家向代理委派工作:科学家设定研究目标与方法论指导,代理负责执行所有分析流程。每个任务包含研究目标、方法指引及源自已发表研究的原始数据,代理需完成一系列分析以达成目标。分析产出(如峰调用矩阵、差异表达表)由程序化方式比对原研究结果进行评分。在涵盖138项独立数据产物的20个任务中,13个前沿模型得分介于0.00(Gemini 3.1 Flash Lite)至0.48(GPT 5.6 Sol)之间。代理在超100 GB数据任务中得分仅为0.10,低于<100 GB任务的0.36;在需3步以上分析的任务中,平均分降至0.24。平均完成时间6.8小时,消耗10200万次令牌,花费43美元,最长尝试耗时24小时、使用10.7亿令牌、支出525美元。值得注意的是,得分最高的代理反而更节省资源。结果表明,大语言模型在执行多步骤分析、管理大规模数据、跨领域协作方面存在显著差异。
原文摘要 · Abstract (English)
Artificial intelligence (AI) promises to accelerate biological research by automating computational analyses. Yet the ability of AI agents to carry out computational biology at the scale of complete research studies has not been systematically evaluated. Here we introduce BixBench3, a benchmark that measures the capacity of AI agents to process raw biological data through to scientific results. We designed BixBench3 tasks to mirror the delegation of work from a scientist to an agent: the scientist chooses the research question and high-level methods, then delegates implementation of all analyses to the agent. In each task, an agent receives a research objective, methodological guidance, and raw data derived from a published scientific study, and must execute a sequence of analyses to achieve the research objective. The data artifacts resulting from these analyses - such as peak call matrices or differential expression tables - are programmatically graded against the corresponding artifacts generated and reported in the original study. Across 20 BixBench3 tasks encompassing the generation of 138 unique artifacts, we find that 13 frontier models achieve scores ranging from 0.00 for Gemini 3.1 Flash Lite to 0.48 for GPT 5.6 Sol. Agents perform worse on tasks with larger raw datasets (0.36 on tasks with <100 GB versus 0.10 on tasks with >100 GB) and on analyses requiring more sequential steps (0.36 at 1-2 steps vs 0.24 at 3+). On average, agents use 6.8 hours, 102 million tokens, and $43 to complete each task, with the longest attempts consuming 24 hours, 1.07 billion tokens, and $525. Notably, the highest-scoring agents used fewer tokens and were cheaper than less performant options. These results reveal that LLMs vary substantially in their ability to (1) execute multiple sequential analysis steps coherently, (2) manage large quantities of raw data, and (3) work across scientific domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。