arXiv:2603.10395cs.LG2026-03中稿 · ICML被引 3

用强化学习训练图生成模型,让生成的分子更优更高效。

Graph-GRPO: Training Graph Flow Models with Reinforcement Learning

  • 通过解析推导实现可微分的图生成过程,替代传统采样方法。
  • 仅用50步去噪即在平面和树结构上达95%以上有效新颖度。
  • 适合需要优化分子结构的任务,如新药研发中的分子设计。

图生成是药物发现等广泛应用的基础任务。近年来,基于离散流匹配的图流模型(GFM)因其优异性能和灵活采样而受到关注。然而,如何有效对齐GFM与复杂人类偏好或特定任务目标仍是挑战。本文提出Graph-GRPO,一种基于可验证奖励的在线强化学习框架来训练GFM。主要贡献:(1)推导出GFM的转移概率解析表达式,取代蒙特卡洛采样,实现完全可微的轨迹演进;(2)提出一种局部扰动与重生成策略,随机调整图中特定节点和边并重新生成,支持局部探索与自我优化。在合成与真实数据集上的大量实验表明,仅需50步去噪,该方法在平面图和树图数据集上分别取得95.0%和97.5%的Valid-Unique-Novelty得分。此外,在分子优化任务中,Graph-GRPO优于基于图与片段的强化学习方法及经典遗传算法,达到当前最佳水平。

原文摘要 · Abstract (English)

Graph generation is a fundamental task with broad applications, such as drug discovery. Recently, discrete flow matching-based graph generation, \aka, graph flow model (GFM), has emerged due to its superior performance and flexible sampling. However, effectively aligning GFMs with complex human preferences or task-specific objectives remains a significant challenge. In this paper, we propose Graph-GRPO, an online reinforcement learning (RL) framework for training GFMs under verifiable rewards. Our method makes two key contributions: (1) We derive an analytical expression for the transition probability of GFMs, replacing the Monte Carlo sampling and enabling fully differentiable rollouts for RL training; (2) We propose a refinement strategy that randomly perturbs specific nodes and edges in a graph, and regenerates them, allowing for localized exploration and self-improvement of generation quality. Extensive experiments on both synthetic and real datasets demonstrate the effectiveness of Graph-GRPO. With only 50 denoising steps, our method achieves 95.0\% and 97.5\% Valid-Unique-Novelty scores on the planar and tree datasets, respectively. Moreover, Graph-GRPO achieves state-of-the-art performance on the molecular optimization tasks, outperforming graph-based and fragment-based RL methods as well as classic genetic algorithms.

图生成强化学习分子设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。