首个动态4D药物-蛋白解离数据集,助力AI预测药物代谢速率。
A Novel 4-D Dataset Paradigm for Studying Complete Ligand-Protein Dissociation Dynamics
- 构建4D动态轨迹数据库DD-13M,涵盖26,000次完整解离过程。
- 训练得到UnbindingFlow模型,可生成新靶点解离轨迹并预测koff值。
- 适合药物研发人员研究药物代谢动力学,推动AI驱动的药物设计。
药物-蛋白结合与解离的动力学和动态过程对理解药物吸收与代谢至关重要。尽管人工智能(AI)工具在药物-蛋白相互作用研究中取得进展,现有训练数据集仍局限于静态结构或准静态构象。本文提出一种新型计算方法,可快速生成药物-蛋白解离轨迹,并发布首个时间分辨的4-D(t, x, y, z)轨迹数据库DD-13M。该数据集捕获了565个配体-蛋白复合物的超过26,000次完整解离过程,提供近1300万帧全原子模拟轨迹。基于此数据集,训练了一个深度等变生成模型UnbindingFlow,可为新靶点生成解离轨迹并准确预测其解离速率常数(koff)。DD-13M为AI模型提供了全新类型的训练数据,建立了一种研究药物-蛋白相互作用动态的新范式。
原文摘要 · Abstract (English)
The kinetics and dynamics of drug-protein binding and dissociation are crucial to understanding drug absorption and metabolism. Despite advances in artificial intelligence (AI) tools for drug-protein interaction studies, existing training datasets remain limited to static structures or quasi-static conformations. This paper proposes a novel computational approach for rapidly generating drug-protein dissociation trajectories and presents the inaugural dynamically time-resolved 4-D (t, x, y, z) trajectory database DD-13M. This dataset captures over 26,000 complete dissociation processes for 565 ligand-protein complexes, providing nearly 13 million frames of all-atom simulation trajectories. A deep equivariant generative model, UnbindingFlow, was trained using the DD-13M dataset. This model has the capacity to produce dissociation trajectories for novel targets whilst accurately predicting their rate constants (koff). DD-13M introduces a new type of training dataset for AI models, establishing a de novo paradigm for studying the dynamics of drug-protein interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。