用嵌套蒙特卡洛方法改进文本生成的推理期控制,更准更稳。
Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo

- 引入嵌套序贯蒙特卡洛,解决采样偏差和权重退化问题
- 在毒性与流畅性任务中,性能优于传统方法
- 适合需要精准控制生成质量的研究者和工程师
我们研究离散扩散语言模型中的推理期控制,目标是在不重新训练的前提下,引导采样过程趋向序列级奖励。以往工作主要依赖粒子方法,如 best-of-$n$ 采样和自举序贯蒙特卡洛,分别存在过度乐观和权重退化的问题。本文提出用于费曼-卡茨导向的嵌套序贯蒙特卡洛(NSMC)与全自适应嵌套序贯蒙特卡洛(FA-NSMC),识别并修正了先前公式中的误差,从而获得无偏最终估计。在毒性与流畅性控制任务上的评估表明,NSMC 和 FA-NSMC 均持续优于 best-of-$n$ 与自举 SMC。
原文摘要 · Abstract (English)
We study inference-time control for text generation in discrete diffusion language models, where the goal is to steer sampling toward sequence-level rewards without retraining. Prior work in this domain has focused on particle-based methods such as best-of-$n$ sampling and bootstrap sequential Monte Carlo, which may suffer from overoptimism and weight degeneracy, respectively. We address these limitations using \emph{nested} sequential Monte Carlo methods. We formulate nested SMC (NSMC) and fully-adapted nested SMC (FA-NSMC) for Feynman--Kac steering, identifying and correcting errors in prior formulations that lead to biased final estimates. We evaluate these methods on toxicity and fluency steering tasks, showing that NSMC and FA-NSMC consistently outperform best-of-$n$ and bootstrap SMC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。