用排序模型替代回归模型,提升分子优化效率
Ranking over Regression for Bayesian Optimization and Molecule Selection
- 用排序模型代替传统回归模型做代理预测
- 在结构-性质关系粗糙的数据集上表现更优
- 早期迭代中仍保持高排序能力,适合新分子设计
贝叶斯优化(BO)已成为自动驾驶、药物与材料发现等场景中自主决策的关键工具。随着自驱动实验室兴起,基于机器学习的化学系统优化愈发重要。传统BO使用回归模型预测未知区域的分布,但在分子选择中,属性的相对排序可能比精确数值更重要。本文提出基于排序的贝叶斯优化(RBO),采用排序模型作为代理。我们在多个化学数据集上系统对比了RBO与传统BO的性能,结果表明:在结构-性质关系不连续及存在活性悬崖的数据集中,RBO表现相当或更优;且代理模型的排序能力与优化性能高度相关,即使在优化早期也保持稳定。结论表明,RBO是新型化合物优化的有效替代方案。
原文摘要 · Abstract (English)
Bayesian optimization (BO) has become an indispensable tool for autonomous decision-making across diverse applications from autonomous vehicle control to accelerated drug and materials discovery. With the growing interest in self-driving laboratories, BO of chemical systems is crucial for machine learning (ML) guided experimental planning. Typically, BO employs a regression surrogate model to predict the distribution of unseen parts of the search space. However, for the selection of molecules, picking the top candidates with respect to a distribution, the relative ordering of their properties may be more important than their exact values. In this paper, we introduce Rank-based Bayesian Optimization (RBO), which utilizes a ranking model as the surrogate. We present a comprehensive investigation of RBO's optimization performance compared to conventional BO on various chemical datasets. Our results demonstrate similar or improved optimization performance using ranking models, particularly for datasets with rough structure-property landscapes and activity cliffs. Furthermore, we observe a high correlation between the surrogate ranking ability and BO performance, and this ability is maintained even at early iterations of BO optimization when using ranking surrogate models. We conclude that RBO is an effective alternative to regression-based BO, especially for optimizing novel chemical compounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。