arXiv:2507.04832cs.LG2025-07被引 9

通过分步对齐提升离散扩散模型的生成质量,显著优于强化学习基线。

Discrete Diffusion Trajectory Alignment via Stepwise Decomposition

  • 将轨迹对齐分解为每一步的后验匹配,实现高效优化
  • 在DNA设计中预测活性提升12%,语言模型GSM8K得分达81.2
  • 适用于任意奖励函数,特别适合需精准控制生成过程的任务

离散扩散模型在建模各类序列数据方面展现出巨大潜力,包括人类语言和生物序列。受强化学习在语言模型中成功应用的启发,研究者开始关注通过与特定奖励对齐来进一步提升模型性能。本文提出一种离线偏好优化方法,用于实现离散扩散模型的轨迹对齐。不同于传统在最终输出施加奖励并反向传播梯度的方式,我们通过匹配每一步的后验分布,将问题分解为一系列分步对齐目标。该框架实现了高效的扩散优化,兼容任意奖励函数,并在轨迹奖励满足可加分解条件下,获得等价最优解。在多个领域实验中,包括DNA序列设计、蛋白质逆折叠和语言建模,结果均显示本方法显著优于现有方法。尤其在DNA序列设计中,预测活性相比最强的强化学习基线最高提升12%;在语言建模任务中,LLaDA-8B-Instruct模型的GSM8K得分从78.6提升至81.2。

原文摘要 · Abstract (English)

Discrete diffusion models have demonstrated great promise in modeling various sequence data, ranging from human language to biological sequences. Inspired by the success of RL in language models, there is growing interest in further improving the models by alignment with a certain reward. In this work, we propose an offline preference optimization method to approach trajectory alignment for discrete diffusion models. Instead of applying the reward on the final output and backpropagating the gradient to the entire denoising process, we decompose the problem into a set of stepwise alignment objectives by matching the per-step posterior. This framework enables efficient diffusion optimization, is compatible with arbitrary reward functions, and importantly, yields an equivalent optimal solution under additive factorization of the trajectory reward. Experiments across multiple domains including DNA sequence design, protein inverse folding, and language modeling consistently demonstrate the superiority of our approach. Notably, it achieves an up to 12\% improvement over the most competitive RL-based baseline in terms of predicted activity on DNA sequence design, and further improves the GSM8K score from 78.6 to 81.2 on LLaDA-8B-Instruct for language modeling.

离散扩散轨迹对齐偏好优化序列生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。