首个统一生成蛋白构象与动态轨迹的自回归模型
ConfRover: Simultaneous Modeling of Protein Conformation and Dynamics via Autoregression
- 基于自回归架构,分三模块建模蛋白构象与时间动态
- 在ATLAS数据集上实现构象与轨迹联合采样,支持多种下游任务
- 适合蛋白质动力学研究者、结构生物学与生成模型方向
理解蛋白质动态对揭示其生物功能至关重要。随着分子动力学(MD)数据的增多,深度生成模型可高效探索蛋白质构象空间。然而现有方法或未能显式捕捉构象间的时间依赖性,或不支持直接生成无时间依赖的样本。为此,我们提出ConfRover,一种自回归模型,能从MD轨迹中同时学习蛋白质构象与动态,支持时间相关和无关采样。模型核心为模块化设计:(i) 编码层,源自蛋白质折叠模型,将每帧的蛋白质特异性信息与构象嵌入潜在空间;(ii) 时序模块,序列模型,捕捉跨帧的构象动态;(iii) SE(3)扩散模型作为结构解码器,在连续空间生成构象。在涵盖多样结构的大规模蛋白质MD数据集ATLAS上的实验表明,该模型在学习构象动态及支持多种下游任务方面效果显著。ConfRover是首个在单一框架内同时采样蛋白质构象与轨迹的模型,为从蛋白质MD数据中学习提供了新颖且灵活的方法。
原文摘要 · Abstract (English)
Understanding protein dynamics is critical for elucidating their biological functions. The increasing availability of molecular dynamics (MD) data enables the training of deep generative models to efficiently explore the conformational space of proteins. However, existing approaches either fail to explicitly capture the temporal dependencies between conformations or do not support direct generation of time-independent samples. To address these limitations, we introduce ConfRover, an autoregressive model that simultaneously learns protein conformation and dynamics from MD trajectories, supporting both time-dependent and time-independent sampling. At the core of our model is a modular architecture comprising: (i) an encoding layer, adapted from protein folding models, that embeds protein-specific information and conformation at each time frame into a latent space; (ii) a temporal module, a sequence model that captures conformational dynamics across frames; and (iii) an SE(3) diffusion model as the structure decoder, generating conformations in continuous space. Experiments on ATLAS, a large-scale protein MD dataset of diverse structures, demonstrate the effectiveness of our model in learning conformational dynamics and supporting a wide range of downstream tasks. ConfRover is the first model to sample both protein conformations and trajectories within a single framework, offering a novel and flexible approach for learning from protein MD data. Project website: https://bytedance-seed.github.io/ConfRover.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。