arXiv:2608.26585cs.LGcs.CE2026-08

提出GRAS方法,让离散扩散模型无需训练即可高效优化序列奖励。

GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

论文配图:GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion
图 1 · 摘自论文原文
  • 用方差降低的引导提案和自适应采样温度提升推理阶段奖励对齐效果。
  • 在DNA与蛋白质设计任务中表现优于现有免训练方法,媲美微调模型。
  • 适用于可导与不可导奖励,无需额外计算成本,适合实际部署。

离散扩散模型已成为序列数据生成的主流方法,而在不重新训练的情况下,仅通过推理阶段引导其达到下游奖励目标日益重要。此类免训练引导通常采用梯度指引、搜索或两者结合。本文研究组合策略,发现其存在两个缺陷:引导提案从单一噪声样本估计梯度,搜索阶段则以固定温度重采样粒子,忽略了每一步去噪过程中奖励分布的变化。为此,我们引入一系列低成本改进:对可导奖励采用Rao-Blackwell化揭示降低估计方差,对不可导奖励使用留一法基线;在搜索中将每步值标准化为组内相对优势,并证明其可简化为单一有效成分——自适应重采样温度。所提方法称为引导低方差提案与自适应选择(GRAS)。GRAS简单而高效,在调控性DNA与蛋白质设计任务中实现最优免训练奖励,超越已有免训练方法,达到甚至超过奖励微调模型的表现,且对不可导奖励依然有效。

原文摘要 · Abstract (English)

Discrete diffusion models have become a strong, widely adopted class of generators for sequence data, and steering them toward a downstream reward at inference time, without any retraining, is increasingly important. Such training-free steering is done by gradient guidance, by search, or by combining the two. We study the combined regime and identify two weaknesses in how it is usually run: the guided proposal estimates its gradient from a single noisy sample, and the search then resamples particles at a fixed temperature that ignores how rewards spread across each denoising step. We address both with a small set of changes that add no denoiser cost. For the proposal, we lower the estimator variance with a Rao-Blackwellized reveal for differentiable rewards and a leave-one-out baseline for non-differentiable ones; for the search, we standardize the per-step values into a group-relative advantage and prove it collapses to a single active ingredient, an adaptive resampling temperature. We call the resulting method Guided Reduced-variance proposals and Adaptive Selection (GRAS). GRAS is simple yet effective: across regulatory DNA and protein design it attains the best training-free reward, outperforming prior training-free methods and matching or surpassing a reward-fine-tuned model, and it remains effective even for non-differentiable rewards.

离散扩散免训练序列生成奖励对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。