AMix-1通过测试时扩展实现蛋白质设计的高效进化,性能随验证预算提升。
AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model
- 基于贝叶斯流网络与多重序列比对,实现进化信号感知的上下文学习。
- 成功设计出活性提升50倍的AmeR突变体,验证框架有效性。
- 支持测试时可扩展的算法,适合蛋白工程与自动化实验设计研究者。
我们提出AMix-1,一个基于贝叶斯流网络的蛋白质基础模型,结合系统性训练方法,包括预训练缩放定律、涌现能力分析、上下文学习机制和测试时缩放算法。为确保鲁棒可扩展性,建立了预测缩放定律,并揭示了损失视角下结构理解的渐进涌现,最终构建出17亿参数的强大模型。在此基础上,设计了一种基于多重序列比对(MSA)的上下文学习策略,将蛋白质设计统一为通用框架,使AMix-1能识别MSA中的深层进化信号,持续生成结构与功能一致的蛋白质。该框架成功设计出活性较野生型提高最多50倍的AmeR突变体。进一步,引入一种进化式测试时缩放算法,实现体外定向进化,在验证预算增加时带来显著且可扩展的性能提升,为下一代实验室闭环蛋白质设计奠定基础。
原文摘要 · Abstract (English)
We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, encompassing pretraining scaling laws, emergent capability analysis, in-context learning mechanism, and test-time scaling algorithm. To guarantee robust scalability, we establish a predictive scaling law and reveal the progressive emergence of structural understanding via loss perspective, culminating in a strong 1.7-billion model. Building on this foundation, we devise a multiple sequence alignment (MSA)-based in-context learning strategy to unify protein design into a general framework, where AMix-1 recognizes deep evolutionary signals among MSAs and consistently generates structurally and functionally coherent proteins. This framework enables the successful design of a dramatically improved AmeR variant with an up to $50\times$ activity increase over its wild type. Pushing the boundaries of protein engineering, we further empower AMix-1 with an evolutionary test-time scaling algorithm for in silico directed evolution that delivers substantial, scalable performance gains as verification budgets are intensified, laying the groundwork for next-generation lab-in-the-loop protein design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。