用AI代理让蛋白设计更智能,能处理非常规氨基酸。
Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents
- 用大模型+罗塞塔软件构建可自主优化的蛋白设计代理
- 在常规和非常规氨基酸设计上表现媲美专业模型与专家
- 环境设计比提示工程更重要,是打通AI与科学软件的关键
大型语言模型(LLM)具备推理与工具使用能力,为执行复杂科学任务的自主代理创造了机会。蛋白设计是理想的测试场景:尽管机器学习方法表现良好,但主要局限于标准氨基酸和狭窄目标,缺乏通用设计流程工具。我们提出Agent Rosetta,一个结合结构化环境的LLM代理,可操作领先的基于物理的异聚物设计软件Rosetta,支持非标准构建单元和几何建模。该代理通过迭代优化实现用户定义目标,融合了大模型的推理能力与罗塞塔的通用性。我们在标准氨基酸设计上评估,性能匹配专用模型与专家基准;在非标准残基设计中,传统机器学习方法失效,而Agent Rosetta仍达到可比效果。关键发现是,仅靠提示工程难以生成有效罗塞塔指令,证明环境设计对集成大模型与专业软件至关重要。结果表明,合理设计的环境能使大模型代理高效调用科学软件,性能媲美专用工具与人类专家。
原文摘要 · Abstract (English)
Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute complex scientific tasks. Protein design provides a natural testbed: although machine learning (ML) methods achieve strong results, these are largely restricted to canonical amino acids and narrow objectives, leaving unfilled need for a generalist tool for broad design pipelines. We introduce Agent Rosetta, an LLM agent paired with a structured environment for operating Rosetta, the leading physics-based heteropolymer design software, capable of modeling non-canonical building blocks and geometries. Agent Rosetta iteratively refines designs to achieve user-defined objectives, combining LLM reasoning with Rosetta's generality. We evaluate Agent Rosetta on design with canonical amino acids, matching specialized models and expert baselines, and with non-canonical residues -- where ML approaches fail -- achieving comparable performance. Critically, prompt engineering alone often fails to generate Rosetta actions, demonstrating that environment design is essential for integrating LLM agents with specialized software. Our results show that properly designed environments enable LLM agents to make scientific software accessible while matching specialized tools and human experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。