arXiv:2606.28301cs.LGcs.DS2026-06

提出高效奖励引导采样方法,提升掩码扩散模型生成质量与修复能力

VGB for Masked Diffusion Model: Efficient Test-time Scaling for Reward Satisfaction and Sample Editing

  • 基于回溯随机游走构建可动态重掩码的图结构采样器
  • 在数独和QM9任务上实现高奖励生成,且计算复杂度为二次方
  • 适合需约束满足或样本修复的生成场景,如科学建模与逻辑推理

推理时缩放是提升生成模型性能的有前景范式,尤其在输出需满足结构约束或优化下游奖励时。本文针对掩码扩散模型(MDM),提出MDM-VGB,一种离散扩散采样器,通过理论严谨的奖励引导重掩码机制增强去掩码生成过程。受经典Jerrum-Sinclair回溯马尔可夫链在奖励倾斜生成中成功启发,MDM-VGB将回溯随机游走从固定前缀树扩展至掩码状态图,允许任意位置的解掩码与重掩码。该采样器偏好能带来更高价值部分配置的移动,从而实现高效高奖励生成与低奖励样本的快速修复。理论上证明其对过程-验证器噪声具有鲁棒性,且复杂度为二次方;而流行的测试时启发式如best-of-N可能因误差累积导致指数级复杂度。理论结论在多个主流约束满足与科学基准(如Sudoku、QM9)上得到强实证支持。

原文摘要 · Abstract (English)

Inference-time scaling is a promising paradigm to improve generative models, especially when outputs must satisfy structural constraints or optimize downstream rewards. We consider Masked Diffusion Model (MDM) and introduce MDM-VGB, a discrete diffusion sampler that augments unmasking generation with theoretically principled reward-guided remasking. Inspired by the recent success of the classical Jerrum-Sinclair backtracking Markov chain in reward-tilted generation, MDM-VGB extends the backtracking random walk from a fixed prefix tree to a masked-state graph, allowing tokens to be unmasked and remasked at arbitrary positions. The resulting sampler favors unmasking and remasking moves that lead to higher-value partial configurations, enabling both effective high-reward generation and efficient repair of low-reward samples. We prove that MDM-VGB is robust to process-verifier noise and achieves quadratic complexity, while popular test-time heuristics such as best-of-$N$ can incur exponential complexity due to error accumulation. Our theoretical findings are corroborated by strong empirical performance, particularly on popular constraint-satisfaction and scientific benchmarks such as Sudoku and QM9.

扩散模型奖励引导生成修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。