arXiv:2510.02291cs.LGcs.CV2025-10被引 5

提出新方法提升离散扩散模型的后验采样效果,兼顾速度与控制力。

Test-Time Anchoring for Discrete Diffusion Posterior Sampling

  • 用量化期望实现离散空间的梯度类引导,避免依赖导数
  • 通过锚定重掩码自适应解码,突破维度诅咒限制
  • 适用文本图像生成、风格迁移及大模型问答,无需重新训练

尽管连续扩散模型已取得显著进展,离散扩散为文本与图像联合建模提供统一框架。相比连续模型,离散扩散具备更快推理速度、更精细控制以及无需训练的指导能力,特别适合后验采样。现有离散扩散后验采样方法面临三大挑战:无导数引导导致信号稀疏,连续松弛限制适用性,分裂吉布斯采样受维度诅咒影响。为此,我们提出锚定后验采样(APS),基于两项核心创新:离散嵌入空间中的量化期望以实现类梯度引导,以及锚定重掩码实现自适应解码。APS在标准图像基准上对线性和非线性逆问题均达到当前最优性能。通过训练自由风格化和文本引导编辑验证其通用性,并应用于大规模扩散语言模型,在问答任务中持续提升表现。

原文摘要 · Abstract (English)

While continuous diffusion models have achieved remarkable success, discrete diffusion offers a unified framework for jointly modeling text and images. Beyond unification, discrete diffusion provides faster inference, finer control, and principled training-free guidance, making it well-suited for posterior sampling. Existing approaches to posterior sampling using discrete diffusion face severe challenges: derivative-free guidance yields sparse signals, continuous relaxations limit applicability, and split Gibbs samplers suffer from the curse of dimensionality. To overcome these limitations, we introduce Anchored Posterior Sampling (APS), built on two key innovations: quantized expectation for gradient-like guidance in discrete embedding space, and anchored remasking for adaptive decoding. APS achieves state-of-the-art performance among discrete diffusion samplers on both linear and nonlinear inverse problems across the standard image benchmarks. We demonstrate the generality of APS through training-free stylization and text-guided editing. We further apply APS to a large-scale diffusion language model, showing consistent improvement in question answering.

离散扩散后验采样无训练指导图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。