揭示大模型解码中局部归一化扭曲的本质问题并量化其影响
Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models
- 将主流解码算法视为遍历理论中的平衡态,建立统一理论框架
- 发现局部归一化导致概率分布扭曲,使生成文本质量与多样性下降
- 为解码器设计和机器生成文本检测提供可量化的理论依据
硬件与语言模型架构的进步推动了自然语言生成的革命。然而,自回归模型对下一个词的概率分布进行计算,而从这些分布中采样的过程——即解码——却远未得到足够重视。现有解码策略多依赖启发式方法,难以系统性改进。本文通过遍历理论将主流解码算法表述为平衡态,并明确其优化的目标函数。利用该理论,分析了 top-k、核采样(nucleus)和温度采样中必需的局部归一化步骤的影响。我们指出,局部归一化扭曲是解码策略的根本缺陷,并量化了该扭曲的大小及其对生成文本质量与多样性的数学代理指标的影响。研究结果为解码算法设计和机器生成文本的检测提供了理论指导。
原文摘要 · Abstract (English)
Advances in hardware and language model architecture have spurred a revolution in natural language generation. However, autoregressive models compute probability distributions over next-token choices, and sampling from these distributions, known as decoding, has received significantly less attention than other design choices. Existing decoding strategies are largely based on heuristics, resulting in methods that are difficult to apply or improve in a principled manner. We develop the theory of decoding strategies for language models by expressing popular decoding algorithms as equilibrium states in the language of ergodic theory and stating the objective functions they optimize. Using this, we analyze the effect of the local normalization step required to make probabilities sum to one in top-k, nucleus, and temperature sampling. We argue that local normalization distortion is a fundamental defect of decoding strategies and quantify the size of this distortion and its effect on mathematical proxies for the quality and diversity of generated text. This yields conclusions for the design of decoding algorithms and the detection of machine-generated text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。