arXiv:2609.05818cs.AIcs.CY2026-09被引 5

评测大模型用蛋白设计工具的能力,发现部分模型表现优异但规划仍不稳。

Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

论文配图:Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
图 1 · 摘自论文原文
  • 构建新基准测试大模型使用蛋白设计工具的全流程能力。
  • 15个模型中7个拒绝所有任务,最高得分由Claude Sonnet 4和Gemini 3 Pro达成。
  • 适合生物信息学与AI交叉研究者参考,关注工具调用与策略规划能力。

我们提出ABLE,一个评估大语言模型代理使用生物人工智能模型(如ProteinMPNN和AlphaFold3)进行双重用途蛋白设计工作流的能力的基准。ABLE通过结构检索、序列生成和设计验证等任务评估代理表现。我们测试了15个前沿模型,发现其中7个完全拒绝执行任务,其余模型表现差异显著。Claude Sonnet 4和Gemini 3 Pro在信息检索、工具选择和工具使用方面得分最高。进一步将部分任务表现与专家人类基准对比,结果表明当前大模型虽能显著降低蛋白设计门槛,但在规划、策略生成以及整合生物学知识与工具使用方面仍存在明显不一致性。

原文摘要 · Abstract (English)

We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but remain inconsistent in planning, strategy generation, and integrating biological knowledge with tool use.

蛋白设计大模型评测生物AI工具调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。