用少于40次测试找到10个优质候选药物,提升新靶点药物发现效率
SPADE: Faster Drug Discovery by Learning from Sparse Data

- 基于稀疏数据设计自适应筛选策略,动态优化候选分子选择
- 平均仅需40次测试即可获得10个高质量配体,样本效率提升7%-32%
- 适用于无先验数据的新靶点蛋白,比现有方法快10倍
药物发现旨在寻找能强效且特异性结合靶蛋白的分子(配体)。然而,仅有不到5%的候选配体能通过药物发现的早期门槛。我们亟需在缺乏历史数据的新靶点上有效工作的方法。从零开始,必须迭代地选择并测试候选配体,以最少的测试次数找到足够数量的优质配体。本文提出的SPADE算法引入一种新型配体筛选方法,在平均仅40次测试下即可找到10个高质量配体。在一对一比较中,SPADE在更多蛋白质上优于深度学习和贝叶斯优化方法,样本效率中位提升7%-32%。SPADE在打分速度上也比最接近的竞争对手快10倍。数据集与代码已公开。
原文摘要 · Abstract (English)
Drug discovery seeks molecules (ligands) that bind strongly and selectively to a target protein. However, fewer than 5% of candidate ligands pass the bar for even the early stages of drug discovery. Furthermore, we want methods that work for novel proteins for which we have no prior data. Starting from scratch, we have to iteratively select and test candidate ligands such that we find enough ligands of the desired quality in as few tests as possible. Our proposed algorithm, named SPADE, introduces a novel approach to ligand selection that requires only 40 tests on average to find 10 high-quality ligands. In one-vs-one comparisons, SPADE outperforms deep learning and Bayesian optimization methods on more proteins, achieving median improvements of 7%-32% in sample efficiency. SPADE is also 10x faster than its closest competitor at scoring candidate drugs. Dataset and code is available at https://anonymous.4open.science/r/SPADE_Fast_Drug_Discovery_by_Learning_from_Sparse_Data-F028/README.md
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。