arXiv:2605.04291cs.LG2026-05ACL

用预训练语言模型作能量函数,提升离散文本扩散模型质量。

Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion

  • 以预训练语言模型为能量函数,构建基于格伯动力学的文本生成框架。
  • 在零样本常识推理、数独等任务上表现优于现有扩散模型。
  • 生成文本质量媲美同规模自回归模型,适合需要高质量生成的场景。

我们提出一种基于统计物理格伯动力学的离散扩散语言模型。核心思想是不直接训练具有均匀转移核的离散扩散模型,而是利用预训练因果/掩码语言模型构建能量函数。该能量函数作为稳态分布,显著提升了生成文本的质量。将UL2作为预训练模型引入扩散流程后,我们的模型在性能上超越了先前基于扩散的语言模型,并与同等规模的自回归模型表现相当。此外,在零样本常识推理、数独和斑马谜题等规划与搜索任务中,我们的模型也表现出色,甚至优于或媲美以往的扩散模型及GPT-2风格的自回归模型。

原文摘要 · Abstract (English)

We present a discrete diffusion-based language model using Glauber dynamics from statistical physics. Our main insight is that instead of trying to train a discrete state space diffusion model using Glauber dynamics with a uniform transition kernel as the forward process, one can set up an ``energy function'' based on pretrained causal/masked language models. When viewed as the stationary distribution, this energy function allows us to significantly improve the quality of the generated text. Incorporating UL2 as the pretrained model into our diffusion pipeline, we outperform prior diffusion based LMs and perform competitively with autoregressive models of comparable model sizes. Furthermore, our models are competitive with or outperform prior diffusion models and GPT-2 style auto-regressive models on zero-shot common sense reasoning tasks as well as planning and search tasks like Sudoku and Zebra puzzles.

文本生成扩散模型语言模型能量函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。