arXiv:2602.09424cs.LGq-bio.QM2026-02被引 2

用干净样本链提升分子与生物序列生成的奖励引导效果。

Reward-Guided Discrete Diffusion via Clean-Sample Markov Chain for Molecule and Biological Sequence Design

  • 构建基于马尔可夫链的清洁样本采样器,避免依赖噪声中间奖励。
  • 在多种奖励函数下,生成样本奖励均显著高于现有方法。
  • 适合需要高精度奖励引导的药物设计与序列优化任务。

离散扩散模型近年来成为化学与生物学数据生成的强大工具。在这些领域,目标是生成具有高奖励(如分子类药性)的多样化样本,因此奖励引导至关重要。现有方法多依赖中间奖励进行引导,但因科学领域奖励函数非光滑,导致中间奖励噪声大,性能受限。为此,本文提出清洁样本马尔可夫链(CSMC)采样器,实现无需依赖中间奖励的测试时奖励引导采样,支持局部搜索。CSMC利用梅特罗波利斯-赫斯廷斯算法构建清洁样本的马尔可夫链,使其平稳分布为目标分布;通过依次应用前向与反向扩散过程设计提议分布,使接受概率可计算。在分子与生物序列生成任务上,使用多种奖励函数的实验表明,该方法在所有场景中均持续优于依赖中间奖励的先前方法。

原文摘要 · Abstract (English)

Discrete diffusion models have recently emerged as a powerful class of generative models for chemistry and biology data. In these fields, the goal is to generate various samples with high rewards (e.g., drug-likeness in molecules), making reward-based guidance crucial. Most existing methods are based on guiding the diffusion model using intermediate rewards but tend to underperform since intermediate rewards are noisy due to the non-smooth nature of reward functions used in scientific domains. To address this, we propose Clean-Sample Markov Chain (CSMC) Sampler, a method that performs effective test-time reward-guided sampling for discrete diffusion models, enabling local search without relying on intermediate rewards. CSMC constructs a Markov chain of clean samples using the Metropolis-Hastings algorithm such that its stationary distribution is the target distribution. We design a proposal distribution by sequentially applying the forward and backward diffusion processes, making the acceptance probability tractable. Experiments on molecule and biological sequence generation with various reward functions demonstrate that our method consistently outperforms prior approaches that rely on intermediate rewards.

分子生成扩散模型奖励引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。