arXiv:2604.02718cs.LGcs.CL2026-04被引 10

提出评估扩散语言模型的新方法,避免误判生成质量。

Generative Frontiers: Why Evaluation Matters for Diffusion Language Models

  • 用生成困惑度与熵的分解,揭示现有评估的缺陷。
  • 提出生成前沿(generative frontiers)作为更可靠的评估框架。
  • 适合关注模型评估方法论的研究者和实践者。

扩散语言模型近年来取得显著进展,其生成轨迹的灵活性远超自回归模型。这一灵活性推动了大量新方法研究,通常以 GPT-2 small(1.5亿参数)为起点。然而,这些进展带来了评估方法的新问题。本文讨论当前评估方法的局限性,提出有原则的改进方案。首先分析 OpenWebText 为何成为标准基准,而 LM1B 等替代方案本质意义有限。接着指出对扩散模型仅依赖似然性评估的不足,说明单纯使用生成困惑度会导致无信息结果。我们证明生成困惑度与熵是参考分布 KL 散度的两个组成部分,该分解解释了困惑度对熵的敏感性,并自然引出生成前沿作为评估模型生成质量的合理方法。最后报告在该规模下的实证观察结果。配套博客含互动内容,详见 https://patrickpynadath1.github.io/blog/eval_methodology/。

原文摘要 · Abstract (English)

Diffusion language models have seen exciting recent progress, offering far more flexibility in generative trajectories than autoregressive models. This flexibility has motivated a growing body of research into new approaches to diffusion language modeling, which typically begins at the scale of GPT-2 small (150 million parameters). However, these advances introduce new issues with evaluation methodology. In this technical note, we discuss the limitations of current methodology and propose principled augmentations to ensure reliable comparisons. We first discuss why OpenWebText has become the standard benchmark, and why alternatives such as LM1B are inherently less meaningful. We then discuss the limitations of likelihood evaluations for diffusion models, and explain why relying on generative perplexity alone as a metric can lead to uninformative results. To address this, we show that generative perplexity and entropy are two components of the KL divergence to a reference distribution. This decomposition explains generative perplexity's sensitivity to entropy, and naturally suggests generative frontiers as a principled method for evaluating model generative quality. We conclude with empirical observations on model quality at this scale. We include a blog post with interactive content to illustrate the argument at https://patrickpynadath1.github.io/blog/eval_methodology/.

扩散模型评估方法语言模型生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。