arXiv:2509.21737cs.LGcs.AI2025-09被引 9

用多轮强化学习提升药物分子优化效率,仅用500次测试就达到领先效果。

POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization

  • 让大模型学习完整优化过程,而非孤立步骤。
  • 在单属性任务上成功率84%(基线2.3倍),多属性任务50%。
  • 适合需要高效试错的药物研发团队使用。

药物先导化合物优化需在庞大化学空间中迭代调整分子结构,以提升性质同时保持与原始化合物的相似性。现有方法样本效率低,难以用少量实验评估达成良好性能。大语言模型凭借其上下文学习和指令遵循能力,天然契合此类迭代过程,但现有方法未充分挖掘这一优势,将每一步优化视为独立事件。为此,我们提出POLO(Preference-guided multi-turn Optimization for Lead Optimization),使大模型能从完整优化轨迹中学习。核心是偏好引导策略优化(PGPO)算法,在两个互补层面提取学习信号:轨迹级优化强化成功策略,回合级偏好学习通过排序每条路径中的中间分子提供密集对比反馈。通过双重层次的中间评估学习,POLO 充分利用每次昂贵的奥拉克评估,实现更优样本效率。大量实验表明,POLO 在单属性任务上平均成功率84%(比基线高2.3倍),多属性任务达50%,仅用500次奥拉克评估,显著超越当前最先进水平。

原文摘要 · Abstract (English)

Lead optimization in drug discovery requires efficiently navigating vast chemical space through iterative cycles to enhance molecular properties while preserving structural similarity to the original lead compound. Despite recent advances, traditional optimization methods struggle with sample efficiency-achieving good optimization performance with limited oracle evaluations. Large Language Models (LLMs) provide a promising approach through their in-context learning and instruction following capabilities, which align naturally with these iterative processes. However, existing LLM-based methods fail to leverage this strength, treating each optimization step independently. To address this, we present POLO (Preference-guided multi-turn Optimization for Lead Optimization), which enables LLMs to learn from complete optimization trajectories rather than isolated steps. At its core, POLO introduces Preference-Guided Policy Optimization (PGPO), a novel reinforcement learning algorithm that extracts learning signals at two complementary levels: trajectory-level optimization reinforces successful strategies, while turn-level preference learning provides dense comparative feedback by ranking intermediate molecules within each trajectory. Through this dual-level learning from intermediate evaluation, POLO achieves superior sample efficiency by fully exploiting each costly oracle call. Extensive experiments demonstrate that POLO achieves 84% average success rate on single-property tasks (2.3x better than baselines) and 50% on multi-property tasks using only 500 oracle evaluations, significantly advancing the state-of-the-art in sample-efficient molecular optimization.

药物发现强化学习大模型分子优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。