arXiv:2412.10966cs.LGcs.AI2024-12被引 16

FlowDock首次实现多配体蛋白-配体对接与亲和力预测的统一生成建模。

FlowDock: Geometric Flow Matching for Generative Protein-Ligand Docking and Affinity Prediction

  • 基于条件流匹配的几何生成模型,直接从无配体结构生成结合构象。
  • 在PoseBusters上盲对接成功率51%,超越单序列AlphaFold 3。
  • 支持多配体同时建模,适合虚拟筛选与药物发现场景。

近期提出的蛋白质-配体结构生成模型虽强大,但多数无法同时支持柔性对接与亲和力估计。现有方法中,尚无模型能直接建模多个结合配体,且未在药理相关靶点上进行严格评估,限制了其在药物发现中的应用。本文提出FlowDock,首个基于条件流匹配的深度几何生成模型,可直接将任意数量配体的无配体(apo)结构映射到结合态(holo)结构。此外,FlowDock为每个生成的复合物提供结构置信度评分和结合亲和力值,支持新(多配体)药物靶标的快速虚拟筛选。在著名的PoseBusters基准测试中,仅使用无配体蛋白输入且无需多序列比对信息,其盲对接成功率达51%,优于单序列AlphaFold 3;在更具挑战性的DockGen-E数据集上,性能超越单序列AlphaFold 3,与单序列Chai-1相当。在第16届社区结构预测评估竞赛(CASP16)的配体类别中,针对140个蛋白-配体复合物,FlowDock在亲和力预测任务中位列前五,验证了其学习表征在虚拟筛选中的有效性。源代码、数据及预训练模型已公开于https://github.com/BioinfoMachineLearning/FlowDock。

原文摘要 · Abstract (English)

Powerful generative AI models of protein-ligand structure have recently been proposed, but few of these methods support both flexible protein-ligand docking and affinity estimation. Of those that do, none can directly model multiple binding ligands concurrently or have been rigorously benchmarked on pharmacologically relevant drug targets, hindering their widespread adoption in drug discovery efforts. In this work, we propose FlowDock, the first deep geometric generative model based on conditional flow matching that learns to directly map unbound (apo) structures to their bound (holo) counterparts for an arbitrary number of binding ligands. Furthermore, FlowDock provides predicted structural confidence scores and binding affinity values with each of its generated protein-ligand complex structures, enabling fast virtual screening of new (multi-ligand) drug targets. For the well-known PoseBusters Benchmark dataset, FlowDock outperforms single-sequence AlphaFold 3 with a 51% blind docking success rate using unbound (apo) protein input structures and without any information derived from multiple sequence alignments, and for the challenging new DockGen-E dataset, FlowDock outperforms single-sequence AlphaFold 3 and matches single-sequence Chai-1 for binding pocket generalization. Additionally, in the ligand category of the 16th community-wide Critical Assessment of Techniques for Structure Prediction (CASP16), FlowDock ranked among the top-5 methods for pharmacological binding affinity estimation across 140 protein-ligand complexes, demonstrating the efficacy of its learned representations in virtual screening. Source code, data, and pre-trained models are available at https://github.com/BioinfoMachineLearning/FlowDock.

生成模型药物发现对接预测亲和力估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。