用机器学习找蛋白质折叠的最低能量构型,效果比传统方法更好。
Lattice Protein Folding with Variational Annealing
- 用带掩码的RNN和退火机制,高效搜索最优折叠路径。
- 成功预测了60个氨基酸珠子系统的最低能量折叠结构。
- 方法可推广到三维空间和更复杂的蛋白模型,适合生物计算研究者。
理解蛋白质折叠原理是计算生物学的核心,对药物设计、生物工程及基础生命过程认知具有重要意义。晶格蛋白折叠模型提供了一个简化但强大的框架,可用于在约束条件下探索能量最优折叠。然而,寻找这些最优折叠是一个计算上极具挑战性的组合优化问题。本文提出一种新型上界训练方案,通过掩码机制识别二维疏水-极性(HP)晶格蛋白折叠中的最低能量构型。结合膨胀循环神经网络(Dilated RNNs)与温度类波动驱动的退火过程,该方法能准确预测最大含60个氨基酸珠子的基准系统最优折叠。该方案有效屏蔽无效折叠,同时不破坏RNN的自回归采样特性。该方法可推广至三维空间,并适用于更大字母表的晶格蛋白模型。研究结果凸显了先进机器学习技术在解决复杂蛋白质折叠问题及更广泛受限组合优化挑战中的潜力。
原文摘要 · Abstract (English)
Understanding the principles of protein folding is a cornerstone of computational biology, with implications for drug design, bioengineering, and the understanding of fundamental biological processes. Lattice protein folding models offer a simplified yet powerful framework for studying the complexities of protein folding, enabling the exploration of energetically optimal folds under constrained conditions. However, finding these optimal folds is a computationally challenging combinatorial optimization problem. In this work, we introduce a novel upper-bound training scheme that employs masking to identify the lowest-energy folds in two-dimensional Hydrophobic-Polar (HP) lattice protein folding. By leveraging Dilated Recurrent Neural Networks (RNNs) integrated with an annealing process driven by temperature-like fluctuations, our method accurately predicts optimal folds for benchmark systems of up to 60 beads. Our approach also effectively masks invalid folds from being sampled without compromising the autoregressive sampling properties of RNNs. This scheme is generalizable to three spatial dimensions and can be extended to lattice protein models with larger alphabets. Our findings emphasize the potential of advanced machine learning techniques in tackling complex protein folding problems and a broader class of constrained combinatorial optimization challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。