评测生物信息学AI代理的性能与鲁棒性,助力可靠本地化工具开发。
BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
- 构建端到端任务套件,含特定提示与可验证输出。
- 前沿大模型可无额外框架完成多步骤分析,但对干扰敏感。
- 适合关注生物信息学自动化与数据安全的研究者。
我们提出BioAgent Bench,一个用于评估AI代理在常见生物信息学任务中性能与鲁棒性的测评套件。该套件包含人工精心设计的端到端任务(如RNA-seq、变异检测、宏基因组分析),并配有任务专用提示和具体输出成果,支持自动化评估。我们在多个代理框架上测试了前沿闭源与开源模型,并采用基于LLM的评分器判断流程进展与结果有效性。结果显示,基于前沿大模型的代理可在无需复杂定制结构的情况下完成多步生物信息学流程,通常能可靠生成目标最终产物。然而,鲁棒性测试表明,在输入损坏、伪文件干扰及提示冗余等受控扰动下存在失败模式,说明高层流程构建正确并不保证底层推理可靠。通过开源代码及配套资源,我们的主要目标是加速低成本且可靠的本地化代理发展,使其能处理常涉及患者敏感数据或未公开知识产权的复杂生物信息学工作流。
原文摘要 · Abstract (English)
We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatics tasks. The suite consists of manually curated end-to-end tasks (e.g., RNA-seq, variant calling, metagenomics) accompanied by task-specific prompts and concrete output artifacts to support automated assessment. We evaluate frontier closed- and open-weight models across multiple agent harnesses, and use an LLM-based grader to score pipeline progress and outcome validity. We find that agents based on frontier LLMs can complete multi-step bioinformatics pipelines without elaborate custom scaffolding, often producing the requested final artifacts reliably. However, robustness tests reveal failure modes under controlled perturbations (corrupted inputs, decoy files, and prompt bloat), indicating that correct high-level pipeline construction does not guarantee reliable step-level reasoning. By releasing the code and the complementary resources constituting our suite, our primary goal is to accelerate the development of cost-effective yet reliable local agents, capable of handling complex bioinformatics workflows often involving sensitive patient data or unpublished intellectual property.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。