arXiv:2606.01220cs.LGcs.AI2026-06

用强化学习优化扩散模型生成符合靶点结构的药物分子。

Fine-Tuning Diffusion Models for Molecular Generation via Reinforcement Learning and Fast Sampling

论文配图:Fine-Tuning Diffusion Models for Molecular Generation via Reinforcement Learning and Fast Sampling
图 1 · 摘自论文原文
  • 基于强化学习微调扩散模型,实现结构约束下的分子生成。
  • 仅需10步去噪即可快速采样,速度提升显著且保持质量。
  • 适合需要高效生成高多样性药物分子的研究者使用。

在基于结构的药物设计中,同时满足类药性并符合靶蛋白三维结构的分子生成是一个核心挑战。现有生成方法常依赖昂贵的后处理或精心构建的数据集,且在多目标场景下表现有限。为此,本文提出FTDiff,一种针对扩散模型的强化学习微调框架,专用于结构约束下的分子生成。为实现稳定高效的优化,FTDiff采用组相对策略优化(GRPO)策略,并基于无时间预训练扩散模型,引入快速采样机制,将去噪步骤减少至10步,显著加速训练与推理过程,同时保持生成质量。通过优化阈值感知奖励函数,FTDiff有效引导模型生成有效、多样且高质量的分子,平衡多个药物设计目标。在基准数据集上的大量实验表明,FTDiff在不依赖昂贵后处理或复杂数据工程的情况下,持续优于现有方法。

原文摘要 · Abstract (English)

Generating molecules that simultaneously satisfy drug-like properties and conform to the 3D structure of a target protein is a core challenge in structure-based drug design (SBDD). Existing generative approaches, however, often rely on costly post-hoc processing during Sampling or require carefully curated datasets during training, yet still achieve modest gains. These limitations are especially pronounced in multi-objective settings, where balancing conflicting criteria remains a core challenge. To address these challenges, We propose FTDiff, a reinforcement learning fine-tuning framework tailored for diffusion-based molecular generation under structural constraints. To ensure stable and sample-efficient optimization, FTDiff adopts a group relative policy optimization (GRPO) style strategy. Furthermore, FTDiff builds upon a time-free pretrained diffusion model and incorporates a fast sampling mechanism that reduces the number of denoising steps, significantly accelerating both training and inference while maintaining generation quality. By optimizing a fixed threshold-aware reward, FTDiff effectively guides the model to produce valid, diverse, and high- quality molecules that balance multiple drug design objectives. Extensive experiments on benchmark datasets demonstrate that FTDiff consistently outperforms prior methods, without requiring expensive post-hoc optimization or intricate data engineering.

分子生成扩散模型强化学习药物设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。