用压缩的蛋白语言模型嵌入加速蛋白质序列设计,一步生成高质量结果。
ProtFlow: Fast Protein Sequence Design via Flow Matching on Compressed Protein Language Model Embeddings
- 基于压缩的蛋白语言模型嵌入,通过流匹配实现快速生成
- 单步生成即达高质量,推理效率显著高于传统方法
- 适合需要快速设计多链蛋白或抗菌肽的研究者使用
蛋白质序列的设计是蛋白质工程中的核心任务。深度生成方法如自回归模型和扩散模型虽加速了新序列发现,但多关注局部语义,存在推理效率低、建模空间大、训练成本高等问题。为此,我们提出ProtFlow,一种基于流匹配的快速蛋白质序列设计框架,其在蛋白语言模型的语义有意义的隐空间嵌入上操作。通过压缩与平滑隐空间,ProtFlow在有限计算资源下仍保持高性能。借助重流(reflow)技术,实现高质量单步序列生成。此外,我们构建了多链蛋白联合设计流程。在多种任务中评估,包括通用肽段、长链蛋白、抗菌肽和抗体,实验表明ProtFlow优于现有专用方法,展现出在计算蛋白质序列设计与分析中的广泛潜力。
原文摘要 · Abstract (English)
The design of protein sequences with desired functionalities is a fundamental task in protein engineering. Deep generative methods, such as autoregressive models and diffusion models, have greatly accelerated the discovery of novel protein sequences. However, these methods mainly focus on local or shallow residual semantics and suffer from low inference efficiency, large modeling space and high training cost. To address these challenges, we introduce ProtFlow, a fast flow matching-based protein sequence design framework that operates on embeddings derived from semantically meaningful latent space of protein language models. By compressing and smoothing the latent space, ProtFlow enhances performance while training on limited computational resources. Leveraging reflow techniques, ProtFlow enables high-quality single-step sequence generation. Additionally, we develop a joint design pipeline for the design scene of multichain proteins. We evaluate ProtFlow across diverse protein design tasks, including general peptides and long-chain proteins, antimicrobial peptides, and antibodies. Experimental results demonstrate that ProtFlow outperforms task-specific methods in these applications, underscoring its potential and broad applicability in computational protein sequence design and analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。