用主动学习提升分子生成模型的外推能力,让新分子更稳定可合成。
Active Learning Enables Extrapolation in Molecular Generative Models
- 构建闭环主动学习系统,通过量子模拟反馈迭代优化生成模型。
- 生成分子性能超出训练数据范围0.44个标准差,分布外分类准确率提升79%。
- 结合热力学稳定性数据,生成分子稳定性提升3.5倍,适合药物研发者。
尽管生成模型在发现具备理想性质的分子方面前景广阔,但往往难以提出可合成且优于训练数据中已知分子的新分子。我们发现关键瓶颈并非生成过程本身,而是分子性质预测器的泛化能力不足。为此,我们提出一种主动学习驱动的闭环分子生成流程,通过量子化学模拟反馈不断优化生成模型,增强其对新化学空间的泛化能力。相比其他生成模型方法,仅我们的主动学习方案能生成超越训练数据范围的分子(最高达训练数据范围外0.44个标准差),且分布外分子分类准确率提升79%。通过在主动学习循环中引入热力学稳定性数据进行条件生成,所生成分子的稳定性比例是次优模型的3.5倍。
原文摘要 · Abstract (English)
Although generative models hold promise for discovering molecules with optimized desired properties, they often fail to suggest synthesizable molecules that improve upon the known molecules seen in training. We find that a key limitation is not in the molecule generation process itself, but in the poor generalization capabilities of molecular property predictors. We tackle this challenge by creating an active-learning, closed-loop molecule generation pipeline, whereby molecular generative models are iteratively refined on feedback from quantum chemical simulations to improve generalization to new chemical space. Compared against other generative model approaches, only our active learning approach generates molecules with properties that extrapolate beyond the training data (reaching up to 0.44 standard deviations beyond the training data range) and out-of-distribution molecule classification accuracy is improved by 79%. By conditioning molecular generation on thermodynamic stability data from the active-learning loop, the proportion of stable molecules generated is 3.5x higher than the next-best model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。