轻量级蛋白质结构预测模型,效率提升90%以上。
Protenix-Mini+: efficient structure prediction model with scalable pairformer
- 压缩非可扩展操作,降低时间复杂度至线性
- 减少冗余模块,推理延迟下降90%以上
- 适合大分子复合物实时部署,适合资源受限场景
轻量化推理对生物分子结构预测及下游任务至关重要,支持实际应用中的高效部署与推理时扩展。尽管AF3及其变体(如Protenix、Chai-1)已显著提升预测精度,但存在高推理延迟和随序列长度呈立方级增长的时间复杂度,限制了其在大型生物分子复合物中的可扩展性。为此,本文提出三项关键创新:(1) 压缩非可扩展运算以缓解立方时间复杂度;(2) 移除模块间冗余结构以降低计算开销;(3) 采用少步采样器加速原子扩散模块。基于上述设计,我们构建了Protenix-Mini+,一个高度轻量且可扩展的Protenix模型变体。在可接受的性能损失范围内,计算效率显著提升:例如,在低同源性单链蛋白上,Protenix-Mini+相较于完整Protenix模型,内蛋白LDDT下降约3%,但计算效率提升超90%。
原文摘要 · Abstract (English)
Lightweight inference is critical for biomolecular structure prediction and downstream tasks, enabling efficient real-world deployment and inference-time scaling for large-scale applications. While AF3 and its variants (e.g., Protenix, Chai-1) have advanced structure prediction results, they suffer from critical limitations: high inference latency and cubic time complexity with respect to token count, both of which restrict scalability for large biomolecular complexes. To address the core challenge of balancing model efficiency and prediction accuracy, we introduce three key innovations: (1) compressing non-scalable operations to mitigate cubic time complexity, (2) removing redundant blocks across modules to reduce unnecessary overhead, and (3) adopting a few-step sampler for the atom diffusion module to accelerate inference. Building on these design principles, we develop Protenix-Mini+, a highly lightweight and scalable variant of the Protenix model. Within an acceptable range of performance degradation, it substantially improves computational efficiency. For example, in the case of low-homology single-chain proteins, Protenix-Mini+ experiences an intra-protein LDDT drop of approximately 3% relative to the full Protenix model -- an acceptable performance trade-off given its substantially 90%+ improved computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。