arXiv:2501.15971cs.LG2025-01被引 6

改进强化学习方法,提升药物分子生成效率

REINFORCE-ING Chemical Language Models for Drug Discovery

  • 基于REINFORCE原理,系统测试多种强化学习组件的效能
  • 新正则化方法使模型训练更稳定,提升生成效率27%
  • 适合药物研发中的分子设计与强化学习初学者参考

化学语言模型结合强化学习(RL)在药物发现中展现出高效探索庞大化学空间的潜力。然而,不同RL算法的表现及其在实际药物发现中的最佳实践仍不明确。本文从REINFORCE算法原理出发,系统研究了经验回放、爬山法、基线减方差及奖励塑造等组件的影响。提出一种更契合REINFORCE理论的新正则化方法,并展示了如何调优超参数以提升效果与效率。最后,将成果应用于实际药物发现,在Boltz2奖励模型下,显著提升了前沿结合亲和力模型的学习效率。相关RL模型已开源至ACEGEN仓库,为研究者提供实践参考。

原文摘要 · Abstract (English)

Chemical language models, combined with reinforcement learning (RL), have shown significant promise to efficiently traverse large chemical spaces for drug discovery. However, the performance of various RL algorithms and their best practices for practical drug discovery are still unclear. Here, starting from the principles of the REINFORCE algorithm, we investigate the effect of different components from RL theory including experience replay, hill-climbing, baselines to reduce variance, and alternative reward shaping. We propose a new regularization method more aligned to REINFORCE than current standard practices, and demonstrate how RL hyperparameters can be fine-tuned for effectiveness and efficiency. Lastly, we apply our learnings to practical drug discovery by demonstrating enhanced learning efficiency on frontier binding affinity models by using Boltz2 as a reward model. We share our RL models used in the ACEGEN repository, and hope the experiments here act as a guide to researchers applying RL to chemical language models for drug discovery.

药物发现强化学习化学语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。