用预训练联合预测加速分子设计的批量贝叶斯优化。
Pretrained Joint Predictions for Scalable Batch Bayesian Optimization of Molecular Designs
- 基于认知神经网络构建可扩展的结合亲和力联合预测分布。
- 在半合成与真实小分子库上分别减少5倍和10倍迭代次数。
- 适合大规模药物发现中的高通量分子优化场景。
批量合成与测试分子设计是药物研发的关键瓶颈。近年来,利用生物分子基础模型作为代理模型以加速该过程受到广泛关注。本文提出如何构建可扩展的概率代理模型,用于批量贝叶斯优化(Batch BO)。这需要并行的采集函数来权衡不同设计,并能快速从联合预测分布中采样以近似这些函数。通过认知神经网络(ENNs)框架,在结构感知大模型提取的表示基础上,我们获得了结合亲和力的可扩展联合预测分布。本工作关键在于研究了ENNs中先验网络的重要性,并探索了在合成数据上预训练它们以提升下游批量贝叶斯优化性能的方法。实验表明,在半合成基准上可减少最多5倍迭代次数以重新发现已知强效EGFR抑制剂;在真实世界小分子库中,亦可实现最多10倍迭代减少,为大规模药物发现提供了有前景的解决方案。
原文摘要 · Abstract (English)
Batched synthesis and testing of molecular designs is the key bottleneck of drug development. There has been great interest in leveraging biomolecular foundation models as surrogates to accelerate this process. In this work, we show how to obtain scalable probabilistic surrogates of binding affinity for use in Batch Bayesian Optimization (Batch BO). This demands parallel acquisition functions that hedge between designs and the ability to rapidly sample from a joint predictive density to approximate them. Through the framework of Epistemic Neural Networks (ENNs), we obtain scalable joint predictive distributions of binding affinity on top of representations taken from large structure-informed models. Key to this work is an investigation into the importance of prior networks in ENNs and how to pretrain them on synthetic data to improve downstream performance in Batch BO. Their utility is demonstrated by rediscovering known potent EGFR inhibitors on a semi-synthetic benchmark in up to 5x fewer iterations, as well as potent inhibitors from a real-world small-molecule library in up to 10x fewer iterations, offering a promising solution for large-scale drug discovery applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。