arXiv:2410.10431cs.LGq-bio.BM2024-10IJCAI被引 7

用自适应奖励机制提升药物生成多样性,避免陷入局部最优。

Diversity-Aware Reinforcement Learning for de novo Drug Design

  • 引入内在动机与外在奖励结合的动态更新策略
  • 结构与预测双驱动方法显著提升分子多样性
  • 适合需要多样化候选药的药物研发团队

微调预训练生成模型在生成有前景的药物分子方面表现良好。该任务常被建模为强化学习问题,以往方法能有效优化奖励函数以生成潜在药物分子。然而,由于缺乏对奖励函数的自适应更新机制,优化过程易陷入局部最优。局部最优分子的效力未必能在后续药物优化中体现,或作为独立临床候选物有效。因此,生成一组多样化的优质分子至关重要。先前工作通过惩罚结构相似分子来修改奖励函数,主要关注提升奖励值。迄今为止,尚无研究系统探讨不同奖励函数自适应更新机制对生成分子多样性的影响。本文研究了多种内在动机方法及外在奖励惩罚策略,分析其对生成分子多样性的影响。实验表明,结合结构与预测方法通常能获得更好的多样性效果。

原文摘要 · Abstract (English)

Fine-tuning a pre-trained generative model has demonstrated good performance in generating promising drug molecules. The fine-tuning task is often formulated as a reinforcement learning problem, where previous methods efficiently learn to optimize a reward function to generate potential drug molecules. Nevertheless, in the absence of an adaptive update mechanism for the reward function, the optimization process can become stuck in local optima. The efficacy of the optimal molecule in a local optimization may not translate to usefulness in the subsequent drug optimization process or as a potential standalone clinical candidate. Therefore, it is important to generate a diverse set of promising molecules. Prior work has modified the reward function by penalizing structurally similar molecules, primarily focusing on finding molecules with higher rewards. To date, no study has comprehensively examined how different adaptive update mechanisms for the reward function influence the diversity of generated molecules. In this work, we investigate a wide range of intrinsic motivation methods and strategies to penalize the extrinsic reward, and how they affect the diversity of the set of generated molecules. Our experiments reveal that combining structure- and prediction-based methods generally yields better results in terms of diversity.

药物设计强化学习分子生成多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。