arXiv:2511.19423q-bio.QMcs.AI2025-11被引 2

用AI代理系统加速酶设计,自动发现影响氧化还原性质的氨基酸位点。

Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design

  • 构建多工具协同的AI代理框架,融合文献检索、结构解析与电势计算。
  • 在铁硫簇附近定位关键残基突变,实现红移性质预测且效率远超人工。
  • 适合计算生物学家和蛋白质工程研究者,推动自动化科学发现。

我们提出Genie-CAT,一种面向蛋白质设计的工具增强型大语言模型系统,旨在加速科学假说生成。以金属蛋白(如铁氧还蛋白)为例,该系统整合四项能力:基于文献的检索增强生成(RAG)、PDB文件结构解析、静电势计算及红氧化学性质的机器学习预测,形成统一的智能体工作流。通过结合自然语言推理与数据驱动、物理基础计算,系统可生成机制可解释、可验证的序列-结构-功能关联假说。在概念验证中,Genie-CAT自主识别[Fe-S]簇附近影响氧化还原性质的残基级修饰,其结果与专家假设高度一致,耗时仅为传统方法的极小部分。该框架表明,融合语言模型与领域专用工具的AI代理,可弥合符号推理与数值模拟的鸿沟,使大模型从对话助手转变为计算发现的协作伙伴。

原文摘要 · Abstract (English)

We present Genie-CAT, a tool-augmented large-language-model (LLM) system designed to accelerate scientific hypothesis generation in protein design. Using metalloproteins (e.g., ferredoxins) as a case study, Genie-CAT integrates four capabilities -- literature-grounded reasoning through retrieval-augmented generation (RAG), structural parsing of Protein Data Bank files, electrostatic potential calculations, and machine-learning prediction of redox properties -- into a unified agentic workflow. By coupling natural-language reasoning with data-driven and physics-based computation, the system generates mechanistically interpretable, testable hypotheses linking sequence, structure, and function. In proof-of-concept demonstrations, Genie-CAT autonomously identifies residue-level modifications near [Fe--S] clusters that affect redox tuning, reproducing expert-derived hypotheses in a fraction of the time. The framework highlights how AI agents combining language models with domain-specific tools can bridge symbolic reasoning and numerical simulation, transforming LLMs from conversational assistants into partners for computational discovery.

蛋白质设计AI代理红氧化学大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。