用改进采样方法提升生成模型的奖励对齐效率
Psi-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models
- 基于pCNL算法从奖励感知后验初始化粒子,提升采样质量
- 在布局生成、数量感知生成等任务中性能显著优于传统方法
- 适合需要高效奖励对齐的生成模型研究者使用
我们提出Ψ-Sampler,一种基于SMC的推理时奖励对齐框架,采用pCNL算法进行初始粒子采样,以实现与得分模型的有效对齐。近期,基于得分模型的推理时奖励对齐受到广泛关注,标志着从预训练向后训练优化的范式转变。该趋势的核心在于将序列蒙特卡洛(SMC)应用于去噪过程。然而,现有方法通常从高斯先验初始化粒子,难以捕捉与奖励相关的区域,导致采样效率低下。我们证明,从奖励感知后验初始化能显著提升对齐性能。为在高维隐空间中实现后验采样,引入预条件裂解-尼科尔森-朗之万(pCNL)算法,结合维度鲁棒性提议与梯度信息驱动的动力学。该方法实现了高效且可扩展的后验采样,在布局到图像生成、数量感知生成和审美偏好生成等任务中均表现优异。
原文摘要 · Abstract (English)
We introduce $Ψ$-Sampler, an SMC-based framework incorporating pCNL-based initial particle sampling for effective inference-time reward alignment with a score-based generative model. Inference-time reward alignment with score-based generative models has recently gained significant traction, following a broader paradigm shift from pre-training to post-training optimization. At the core of this trend is the application of Sequential Monte Carlo (SMC) to the denoising process. However, existing methods typically initialize particles from the Gaussian prior, which inadequately captures reward-relevant regions and results in reduced sampling efficiency. We demonstrate that initializing from the reward-aware posterior significantly improves alignment performance. To enable posterior sampling in high-dimensional latent spaces, we introduce the preconditioned Crank-Nicolson Langevin (pCNL) algorithm, which combines dimension-robust proposals with gradient-informed dynamics. This approach enables efficient and scalable posterior sampling and consistently improves performance across various reward alignment tasks, including layout-to-image generation, quantity-aware generation, and aesthetic-preference generation, as demonstrated in our experiments. Project Webpage: https://psi-sampler.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。