构建动态蛋白质复合物数据集,助力AI预测未知蛋白相互作用结构。
DynaPPI: A Large-scale Dynamic Protein Dataset for AI-driven Advances in Protein Interactomics

- 基于分子动力学模拟,生成多聚体蛋白从解离到结合的动态轨迹。
- 首次提供涵盖复合物形成全过程的动态数据,填补静态数据空白。
- 为扩散模型训练提供真实动态路径,适合结构生物学与AI交叉研究者。
扩散模型在蛋白主链生成中展现出强大能力,但当前人工智能驱动的生物研究仍无法有效预测未知多链蛋白聚集体(即生物学中的“复合物”)结构。原因在于现有静态或动态蛋白数据集仅关注静态快照或单体轨迹,忽略了多亚基形成复合物的动态过程。为解决这一难题,我们提出了DynaPPI,一个包含蛋白复合物从解离链到结合态完整形成过程的分子动力学(MD)轨迹的大规模动态蛋白数据集。该数据集作为连接静态结构生物学与分子动态交互本质的关键资源,使扩散模型能显式学习已知复合物的动态结合路径,并基于其多样化的生成能力准确预测未知复合物的结构,进一步推动人工智能驱动的结构生物学与蛋白质互作组学发展。
原文摘要 · Abstract (English)
Diffusion models have been widely explored in protein backbone generation due to their powerful generation capabilities.However, in today's AI-driven biological research, predicting the structure of unknown multi-chain protein aggregates (called "complexes" in biology) remains an unsolved challenge.This is because existing static or dynamic protein datasets focus solely on static snapshots or single-entity trajectories, neglecting the dynamic process of multiple monomers forming complexes.To alleviate this dilemma, we present DynaPPI, a dynamic protein dataset comprising molecular dynamics (MD) trajectories of protein complex formation from dissociated chains to the bound state, as a pivotal resource to bridge the gap between static structural biology and the inherently temporal nature of dynamic molecular interactions.Benefiting from this dataset, diffusion models can explicitly learn the dynamic binding trajectories of known complexes and accurately predict the structures of unknown complexes based on their diverse generative properties, thereby further catalyzing AI-driven structural biology and protein interactomics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。