arXiv:2410.18403q-bio.BMcs.LG2024-10ICLR被引 35

用语言模型生成蛋白构象,效率比传统方法快20-100倍。

Structure Language Models for Protein Conformation Generation

  • 将蛋白结构编码为隐空间,用语言模型建模构象分布。
  • 在多个场景中实现20至100倍的生成速度提升。
  • 适合需要高效探索蛋白构象的药物研发人员。

蛋白质通过多种构象执行生物功能,理解这些构象对药物发现至关重要。传统基于物理的模拟方法常难以采样平衡构象且计算成本高。近年来,深度生成模型因其高效性展现出潜力,但多数依赖3D几何空间中的扩散过程,通常局限于亚稳态附近,运行效率不高。本文提出结构语言建模(SLM)框架,先通过离散变分自编码器将蛋白结构编码至紧凑隐空间,再利用条件语言建模捕捉序列特异的构象分布,实现更高效、可解释的多样构象探索。基于该框架,我们结合多种主流语言模型,并提出ESMDiff——一个从ESM3微调的BERT-like结构语言模型,采用掩码扩散机制。在BPTI平衡动力学、构象变化对以及内在无序蛋白等场景中验证了该方法。SLM显著提升效率,生成多样化构象的速度较现有方法提高20至100倍,为未来研究提供新方向。

原文摘要 · Abstract (English)

Proteins adopt multiple structural conformations to perform their diverse biological functions, and understanding these conformations is crucial for advancing drug discovery. Traditional physics-based simulation methods often struggle with sampling equilibrium conformations and are computationally expensive. Recently, deep generative models have shown promise in generating protein conformations as a more efficient alternative. However, these methods predominantly rely on the diffusion process within a 3D geometric space, which typically centers around the vicinity of metastable states and is often inefficient in terms of runtime. In this paper, we introduce Structure Language Modeling (SLM) as a novel framework for efficient protein conformation generation. Specifically, the protein structures are first encoded into a compact latent space using a discrete variational auto-encoder, followed by conditional language modeling that effectively captures sequence-specific conformation distributions. This enables a more efficient and interpretable exploration of diverse ensemble modes compared to existing methods. Based on this general framework, we instantiate SLM with various popular LM architectures as well as proposing the ESMDiff, a novel BERT-like structure language model fine-tuned from ESM3 with masked diffusion. We verify our approach in various scenarios, including the equilibrium dynamics of BPTI, conformational change pairs, and intrinsically disordered proteins. SLM provides a highly efficient solution, offering a 20-100x speedup than existing methods in generating diverse conformations, shedding light on promising avenues for future research.

蛋白构象生成模型语言模型药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。