将扩散模型成功用于离散数据建模,提升语言与图像生成效果。
Discrete Modeling via Boundary Conditional Diffusion Processes
- 先估计离散边界作为先验,再构建条件扩散过程。
- 在翻译、摘要任务上超越现有扩散语言模型,媲美自回归Transformer。
- 在Cifar-10图像生成中达新最优,支持离散像素与类别数据。
我们提出一种高效且有效的框架,将强大的连续扩散过程扩展至离散建模。以往方法因离散数据与连续建模间的差异而受限,研究发现其主要原因是学习概率轮廓时缺乏离散边界的引导。为此,我们设计两步前向过程:先以边界为先验进行估计,再重缩放前向轨迹,构建边界条件扩散模型;反向过程相应调整,确保学习到的轮廓更精准地还原离散数据。实验表明,该方法在语言建模和离散图像生成任务中均表现优异。在三项翻译任务和一项摘要任务中,优于现有最优连续扩散语言模型,性能可媲美自回归Transformer。在使用离散序数像素时,结果与连续扩散模型相当,并在Cifar-10数据集的分类图像生成任务中达到新最优水平。
原文摘要 · Abstract (English)
We present an novel framework for efficiently and effectively extending the powerful continuous diffusion processes to discrete modeling. Previous approaches have suffered from the discrepancy between discrete data and continuous modeling. Our study reveals that the absence of guidance from discrete boundaries in learning probability contours is one of the main reasons. To address this issue, we propose a two-step forward process that first estimates the boundary as a prior distribution and then rescales the forward trajectory to construct a boundary conditional diffusion model. The reverse process is proportionally adjusted to guarantee that the learned contours yield more precise discrete data. Experimental results indicate that our approach achieves strong performance in both language modeling and discrete image generation tasks. In language modeling, our approach surpasses previous state-of-the-art continuous diffusion language models in three translation tasks and a summarization task, while also demonstrating competitive performance compared to auto-regressive transformers. Moreover, our method achieves comparable results to continuous diffusion models when using discrete ordinal pixels and establishes a new state-of-the-art for categorical image generation on the Cifar-10 dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。