构建首个生物信息学大模型智能体综合评测基准,揭示前沿模型实际能力短板。
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- 设计50+真实生物数据分析场景,包含近300个开放问答任务
- 顶尖模型在开放问答中仅17%准确率,多选题表现接近随机
- 面向希望提升生物计算分析能力的研究者与开发者
大型语言模型(LLMs)及其智能体在加速科学研究方面展现出巨大潜力。现有评测基准正从单纯的知识记忆转向更贴近实际科研的任务,如文献综述和实验设计。生物信息学是有望实现完全自主AI发现的领域,但目前尚无全面的评测基准。为此,我们提出生物信息学基准(BixBench),包含超过50个真实世界生物数据分析场景,以及近300个关联的开放问答问题,用于评估基于LLM的智能体探索生物数据集、执行长序列多步骤分析流程及解读复杂结果的能力。我们使用开源的定制化智能体框架对两款前沿模型(GPT-4o 和 Claude 3.5 Sonnet)进行了评估,发现即使是最先进的模型在开放问答中也仅达到17%的准确率,在多选题设置中表现甚至不优于随机猜测。该基准揭示了当前前沿模型在生物信息学任务中的显著局限性,旨在推动具备严谨生物数据分析能力的智能体发展,加速科学发现进程。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future development continue to evolve from pure recall and rote knowledge tasks, towards more practical work such as literature review and experimental planning. Bioinformatics is a domain where fully autonomous AI-driven discovery may be near, but no extensive benchmarks for measuring progress have been introduced to date. We therefore present the Bioinformatics Benchmark (BixBench), a dataset comprising over 50 real-world scenarios of practical biological data analysis with nearly 300 associated open-answer questions designed to measure the ability of LLM-based agents to explore biological datasets, perform long, multi-step analytical trajectories, and interpret the nuanced results of those analyses. We evaluate the performance of two frontier LLMs (GPT-4o and Claude 3.5 Sonnet) using a custom agent framework we open source. We find that even the latest frontier models only achieve 17% accuracy in the open-answer regime, and no better than random in a multiple-choice setting. By exposing the current limitations of frontier models, we hope BixBench can spur the development of agents capable of conducting rigorous bioinformatic analysis and accelerate scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。