arXiv:2604.16896q-bio.QMcs.AI2026-04ACL被引 1

用反思式工具循环设计满足自然语言要求的蛋白质。

ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design

论文配图:ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design
图 1 · 摘自论文原文
  • 通过多轮反馈循环让大模型规划并修正蛋白质设计
  • 在有限标注下实现语言对齐且折叠性良好,优于基线方法
  • 适合需要自然语言指导的蛋白质工程研究者

设计满足自然语言功能要求的蛋白质是蛋白质工程的核心目标。直接微调通用指令微调的大语言模型作为文本到序列生成器虽可行,但需大量数据和计算资源。在监督信息有限的情况下,大语言模型虽能生成连贯的文本计划,却难以可靠地将其转化为实际序列,存在计划-执行差距。为此,我们提出ProtoCycle,一种以大语言模型为主导的代理式蛋白质设计框架,采用多轮、反馈驱动的决策循环。ProtoCycle将大语言模型规划器与轻量级工具环境结合,模拟人类蛋白质工程的迭代流程,并利用大语言模型对工具反馈进行反思以修订计划。该框架通过监督轨迹和在线强化学习训练,在保持良好折叠性的同时实现强语言对齐;消融实验表明,反思机制显著提升序列质量。

原文摘要 · Abstract (English)

Designing proteins that satisfy natural language functional requirements is a central goal in protein engineering. A straightforward baseline is to fine-tune generic instruction-tuned LLMs as direct text-to-sequence generators, but this is data- and compute-hungry. With limited supervision, LLMs can produce coherent plans in text yet fail to reliably realize them as sequences. This plan-execute gap motivates ProtoCycle, an agentic framework for protein design that uses LLMs primarily to drive a multi-round, feedback-driven decision cycle. ProtoCycle couples an LLM planner with a lightweight tool environment designed to emulate the iterative workflow of human protein engineering and uses LLM-driven reflection on tool feedback to revise plans. Trained with supervised trajectories and online reinforcement learning, ProtoCycle achieves strong language alignment while maintaining competitive foldability, and ablations show that reflection substantially improves sequence quality.

蛋白质设计大模型工具增强反思机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。