arXiv:2511.22311cs.AIcond-mat.mes-hall2025-11被引 6

用多个LLM代理协同设计蛋白质序列,无需微调即可高效生成功能蛋白。

Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation

  • 每个氨基酸位置由独立LLM代理负责,通过迭代反馈优化序列
  • 在数小时内完成设计,实验验证了α螺旋和无规卷曲结构蛋白的有效性
  • 无需预训练数据或特定任务调整,适合广泛生物分子设计场景

从头设计具有特定结构、理化性质和功能的蛋白质仍是生物技术、医学与材料科学中的重大挑战,原因在于序列空间庞大且序列、结构与功能之间存在复杂耦合。现有先进生成方法如蛋白质语言模型(PLMs)和基于扩散的架构,通常需要大量微调、任务专用数据或模型重构,限制了其灵活性与可扩展性。为此,我们提出一种受群体智能启发的去中心化、代理式框架用于从头蛋白质设计。该框架中,多个大型语言模型(LLM)代理并行工作,每个代理负责一个残基位置,通过整合设计目标、局部邻域相互作用以及前序迭代的记忆与反馈,迭代提出上下文感知的突变。这种按位置去中心化的协作机制实现了无需依赖模体骨架或多重序列比对的多样化、明确序列的涌现式设计,并在含α螺旋和无规卷曲结构的蛋白质上通过实验验证。通过残基保守性分析、结构度量指标、序列收敛性及嵌入分析,证明该框架展现出涌现行为,并有效导航蛋白质适应度景观。该方法可在数个GPU小时内实现高效、目标导向的设计,全程无需微调或专门训练,为蛋白质设计提供通用且可适配的解决方案。该方法还可推广至其他生物分子系统及科学发现任务。

原文摘要 · Abstract (English)

Designing proteins de novo with tailored structural, physicochemical, and functional properties remains a grand challenge in biotechnology, medicine, and materials science, due to the vastness of sequence space and the complex coupling between sequence, structure, and function. Current state-of-the-art generative methods, such as protein language models (PLMs) and diffusion-based architectures, often require extensive fine-tuning, task-specific data, or model reconfiguration to support objective-directed design, thereby limiting their flexibility and scalability. To overcome these limitations, we present a decentralized, agent-based framework inspired by swarm intelligence for de novo protein design. In this approach, multiple large language model (LLM) agents operate in parallel, each assigned to a specific residue position. These agents iteratively propose context-aware mutations by integrating design objectives, local neighborhood interactions, and memory and feedback from previous iterations. This position-wise, decentralized coordination enables emergent design of diverse, well-defined sequences without reliance on motif scaffolds or multiple sequence alignments, validated with experiments on proteins with alpha helix and coil structures. Through analyses of residue conservation, structure-based metrics, and sequence convergence and embeddings, we demonstrate that the framework exhibits emergent behaviors and effective navigation of the protein fitness landscape. Our method achieves efficient, objective-directed designs within a few GPU-hours and operates entirely without fine-tuning or specialized training, offering a generalizable and adaptable solution for protein design. Beyond proteins, the approach lays the groundwork for collective LLM-driven design across biomolecular systems and other scientific discovery tasks.

蛋白质设计LLM代理生成模型生物分子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。