通过迭代精炼提升离散扩散模型的测试时扩展效果
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
- 在多尝试元蒙特卡洛框架下,逐轮优化中间状态分布
- 低算力下生成质量显著超越现有最佳方法
- 适合追求高效奖励引导生成的研究者
尽管测试时扩展通过奖励引导生成在离散扩散模型中仍处于探索阶段,但其潜力巨大。本文提出针对离散扩散模型的测试时扩展新方法IterRef,利用奖励引导的去噪-加噪过程,对不一致的中间状态进行迭代精炼。该方法在多尝试元蒙特卡洛(MTM)框架下形式化,证明可收敛至奖励对齐分布。与以往假设当前状态已对齐、仅指导后续转移的方法不同,IterRef显式地在原地精炼每一状态,逐步引导其趋向最优中间分布。在文本与图像多个领域,基于多种离散扩散模型的实验表明,IterRef在奖励引导生成质量上均实现持续提升,尤其在低算力条件下表现突出,远超现有最先进基线。
原文摘要 · Abstract (English)
Test-time scaling through reward-guided generation remains largely unexplored for discrete diffusion models despite its potential as a promising alternative. In this work, we introduce Iterative Reward-Guided Refinement (IterRef), a novel test-time scaling method tailored to discrete diffusion that leverages reward-guided noising-denoising transitions to progressively refine misaligned intermediate states. We formalize this process within a Multiple-Try Metropolis (MTM) framework, proving convergence to the reward-aligned distribution. Unlike prior methods that assume the current state is already aligned with the reward distribution and only guide the subsequent transition, our approach explicitly refines each state in situ, progressively steering it toward the optimal intermediate distribution. Across both text and image domains, we evaluate IterRef on diverse discrete diffusion models and observe consistent improvements in reward-guided generation quality. In particular, IterRef achieves striking gains under low compute budgets, far surpassing prior state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。