用更少实验次数,快速找到优质分子候选。
Sample Efficient Generative Optimization for Molecular Design

- 用生成模型结合贝叶斯优化,智能筛选化学空间。
- 在PMO基准上仅用1/10的评估次数达到顶尖性能。
- 适合药物、材料设计等高成本分子优化场景。
药物发现、材料设计和催化中的分子优化需在巨大化学空间中搜索,但高精度评估或实验测量成本高昂。因此,方法的样本效率——即找到优质候选所需的最少评估次数——至关重要。本文提出样本高效生成优化(SEGO)框架,用于在自动生成的分子上进行贝叶斯优化。在SEGO中,概率代理模型推测潜在优解所在区域,生成模型被引导生成该区域的候选分子,通过采集函数选择最有希望的分子进行评估,其结果既用于更新代理模型,也作为真实奖励锚定生成器。在实际分子优化(PMO)基准测试中,SEGO仅用其他方法十分之一的调用次数即达到最优性能;在多参数对接任务中,其所需评估次数仅为现有方法的一半即可获得十项有效命中。这些提升使分子优化更接近直接基于实验反馈的探索范式。
原文摘要 · Abstract (English)
Molecular optimization in drug discovery, materials design, and catalysis requires searching vast chemical spaces under tight evaluation budgets, since high-fidelity oracles and experimental measurements are costly. The practical impact of an optimization method therefore hinges on its sample efficiency: how few evaluations it needs to find strong candidates. We introduce Sample Efficient Generative Optimization (SEGO), a framework for Bayesian optimization on adaptively generated molecules. In SEGO, a probabilistic surrogate model forms a hypothesis about where hits lie in chemical space, a generative model is steered to propose candidates in that region, the most promising candidate is selected via an acquisition function, and the resulting oracle call is used both to sharpen the surrogate and to anchor the generator in real reward. SEGO attains state-of-the-art performance on the practical molecular optimization (PMO) benchmark using only one tenth of the oracle calls consumed by other methods, and on a multiparameter docking task it reaches ten hits in roughly half the oracle calls of existing approaches. These gains move molecular optimization closer to campaigns driven by direct experimental feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。