arXiv:2504.01389cs.LGphysics.chem-ph2025-04被引 10

用直接偏好优化加速药物分子设计,比现有方法快6%。

De Novo Molecular Design Enabled by Direct Preference Optimization and Curriculum Learning

  • 用分子评分对指导模型生成优质分子,避开传统强化学习难题。
  • 在GuacaMol基准上,帕里那普利多目标任务得分0.883,提升6%。
  • 适合需要高效生成高潜力候选分子的药企与研究机构。

从头分子设计在药物发现和材料科学中有广泛应用。巨大的化学空间使直接搜索计算成本过高,而传统实验筛选耗时耗力。因此,高效的分子生成与筛选方法对加速药物研发、降低成本至关重要。尽管强化学习已用于通过奖励机制优化分子属性,但其实际应用受限于训练效率低、收敛难和稳定性差。为此,我们引入自然语言处理中的直接偏好优化(DPO),利用基于分子评分的样本对,最大化高质量与低质量分子的概率差异,有效引导模型生成更优化合物。此外,结合课程学习进一步提升训练效率并加速收敛。在GuacaMol基准上的系统评估表明,该方法表现优异:例如,在帕里那普利多目标优化任务中得分达0.883,较竞争模型提升6%。后续靶蛋白结合实验也验证了其实际有效性。结果表明,DPO在分子设计任务中具有强大潜力,是一种稳健高效的基于数据驱动的药物发现解决方案。

原文摘要 · Abstract (English)

De novo molecular design has extensive applications in drug discovery and materials science. The vast chemical space renders direct molecular searches computationally prohibitive, while traditional experimental screening is both time- and labor-intensive. Efficient molecular generation and screening methods are therefore essential for accelerating drug discovery and reducing costs. Although reinforcement learning (RL) has been applied to optimize molecular properties via reward mechanisms, its practical utility is limited by issues in training efficiency, convergence, and stability. To address these challenges, we adopt Direct Preference Optimization (DPO) from NLP, which uses molecular score-based sample pairs to maximize the likelihood difference between high- and low-quality molecules, effectively guiding the model toward better compounds. Moreover, integrating curriculum learning further boosts training efficiency and accelerates convergence. A systematic evaluation of the proposed method on the GuacaMol Benchmark yielded excellent scores. For instance, the method achieved a score of 0.883 on the Perindopril MPO task, representing a 6\% improvement over competing models. And subsequent target protein binding experiments confirmed its practical efficacy. These results demonstrate the strong potential of DPO for molecular design tasks and highlight its effectiveness as a robust and efficient solution for data-driven drug discovery.

分子设计强化学习偏好优化药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。