OMTRA统一生成模型,一站式解决药物设计中的多种任务。
OMTRA: A Multi-Task Generative Model for Structure-Based Drug Design
- 基于流匹配的多任务生成框架,统一处理药物设计各类问题。
- 在口袋条件下的从头设计和对接任务中达到顶尖性能。
- 提供5亿个3D分子构象数据集,支持更广化学空间探索。
基于结构的药物设计(SBDD)旨在设计能与特定蛋白质口袋结合的小分子配体。计算方法在现代SBDD流程中至关重要,常通过对接或药效团搜索进行虚拟筛选。近年来,生成建模方法致力于通过从头设计提升新配体发现能力。本文指出,这些任务具有共同结构特征,可视为同一生成建模框架的不同实例。我们提出OMTRA——一种多模态流匹配模型,灵活执行多种与SBDD相关任务,包括传统工作流中无对应的任务。此外,我们构建了一个包含5亿个3D分子构象的数据集,补充蛋白-配体数据,扩展训练可用的化学多样性。OMTRA在口袋条件下的从头设计和对接任务中表现最优;然而,大规模预训练和多任务训练的效果较为有限。所有代码、训练模型及数据集均可在https://github.com/gnina/OMTRA获取。
原文摘要 · Abstract (English)
Structure-based drug design (SBDD) focuses on designing small-molecule ligands that bind to specific protein pockets. Computational methods are integral in modern SBDD workflows and often make use of virtual screening methods via docking or pharmacophore search. Modern generative modeling approaches have focused on improving novel ligand discovery by enabling de novo design. In this work, we recognize that these tasks share a common structure and can therefore be represented as different instantiations of a consistent generative modeling framework. We propose a unified approach in OMTRA, a multi-modal flow matching model that flexibly performs many tasks relevant to SBDD, including some with no analogue in conventional workflows. Additionally, we curate a dataset of 500M 3D molecular conformers, complementing protein-ligand data and expanding the chemical diversity available for training. OMTRA obtains state of the art performance on pocket-conditioned de novo design and docking; however, the effects of large-scale pretraining and multi-task training are modest. All code, trained models, and dataset for reproducing this work are available at https://github.com/gnina/OMTRA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。