用物理反馈强化学习,让扩散模型生成更稳定的分子结构。
Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation
- 用强化学习+物理力场评估,指导扩散模型生成分子
- 在QM9和GEOM-drug上显著提升分子稳定性
- 适合需要高精度分子结构的药物设计研究
生成符合物理规律的3D分子结构仍是分子生成建模的核心挑战。尽管配备等变神经网络的扩散模型在捕捉分子几何方面取得进展,但常难以生成满足力场一致性等物理原则的平衡结构。为此,我们提出基于物理反馈的强化学习框架(RLPF),将3D分子生成建模为马尔可夫决策过程,并采用近端策略优化微调等变扩散模型。关键在于,RLPF引入基于力场评估的奖励函数,提供直接物理反馈,引导生成趋向能量稳定且物理合理的结构。在QM9和GEOM-drug数据集上的实验表明,RLPF显著优于现有方法。结果凸显了在生成模型中融入物理反馈的价值。代码已公开:https://github.com/ZhijianZhou/RLPF/tree/verl_diffusion。
原文摘要 · Abstract (English)
Generating physically realistic 3D molecular structures remains a core challenge in molecular generative modeling. While diffusion models equipped with equivariant neural networks have made progress in capturing molecular geometries, they often struggle to produce equilibrium structures that adhere to physical principles such as force field consistency. To bridge this gap, we propose Reinforcement Learning with Physical Feedback (RLPF), a novel framework that extends Denoising Diffusion Policy Optimization to 3D molecular generation. RLPF formulates the task as a Markov decision process and applies proximal policy optimization to fine-tune equivariant diffusion models. Crucially, RLPF introduces reward functions derived from force-field evaluations, providing direct physical feedback to guide the generation toward energetically stable and physically meaningful structures. Experiments on the QM9 and GEOM-drug datasets demonstrate that RLPF significantly improves molecular stability compared to existing methods. These results highlight the value of incorporating physics-based feedback into generative modeling. The code is available at: https://github.com/ZhijianZhou/RLPF/tree/verl_diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。