用并行马尔可夫链蒙特卡洛加速蛋白质进化参数估计
Boltzmann Machine Learning with a Parallel, Persistent Markov chain Monte Carlo method for Estimating Evolutionary Fields and Couplings from a Protein Multiple Sequence Alignment
- 采用并行持久马尔可夫链采样提高学习效率
- 在8个蛋白家族上实现高精度残基接触预测
- 针对蛋白构象设计超参数调节策略,适合结构生物学研究
从同源蛋白多序列比对中观测到的单位点与成对氨基酸频率出发,求解逆Potts模型以估计蛋白质演化中的单位点场和成对耦合,仍是研究蛋白质结构与进化的重要方法。为确保场和耦合的可重复性,本文采用计算量大的玻尔兹曼机方法。为降低计算耗时,引入并行、持久的马尔可夫链蒙特卡洛方法估算每轮学习中的单点与成对边缘分布,并结合随机梯度下降法进一步提速。另一难点是超参数调节:场与耦合各有两个正则化参数。传统以残基接触预测精度调参不敏感,本文提出基于蛋白构象合理性的条件来调节参数。该方法已在八个蛋白家族上成功应用。
原文摘要 · Abstract (English)
The inverse Potts problem for estimating evolutionary single-site fields and pairwise couplings in homologous protein sequences from their single-site and pairwise amino acid frequencies observed in their multiple sequence alignment would be still one of useful methods in the studies of protein structure and evolution. Since the reproducibility of fields and couplings are the most important, the Boltzmann machine method is employed here, although it is computationally intensive. In order to reduce computational time required for the Boltzmann machine, parallel, persistent Markov chain Monte Carlo method is employed to estimate the single-site and pairwise marginal distributions in each learning step. Also, stochastic gradient descent methods are used to reduce computational time for each learning. Another problem is how to adjust the values of hyperparameters; there are two regularization parameters for evolutionary fields and couplings. The precision of contact residue pair prediction is often used to adjust the hyperparameters. However, it is not sensitive to these regularization parameters. Here, they are adjusted for the fields and couplings to satisfy a specific condition that is appropriate for protein conformations. This method has been applied to eight protein families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。