研究大模型在不同词序下概率一致性,发现其实际表现偏离理论预期。
Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies
- 从理论上证明序列困惑度在任意顺序下应不变
- 实验证明模型在乱序输入时概率分布明显偏离理论值
- 揭示自注意力机制中的位置与局部性偏差是根源
自回归大语言模型在不同词序下能否学习到一致的概率分布?我们严格证明:对于任何定义良好的概率分布,序列困惑度在任意因子分解(包括正序、逆序或任意排列)下保持不变。这一结果为研究大模型如何从数据中学习提供了严谨的理论基础,并定义了规范的实验评估协议。应用该协议后,我们发现以往研究存在关键方法学缺陷。在科学文本上对GPT-2模型进行正序、逆序及任意排列顺序的再训练,结果显示所有顺序均出现系统性偏离理论不变性的现象,其中任意排列顺序的偏差尤为显著,而正序与逆序模型之间则基本一致。偏差可归因于自注意力机制中的位置和局部性偏置。理论与实证结果为理解大模型的位置偏差提供了新路径,并提出检测概率分布不一致性的方法。
原文摘要 · Abstract (English)
Can autoregressive large language models (LLMs) learn consistent probability distributions when trained on sequences in different token orders? We prove formally that for any well-defined probability distribution, sequence perplexity is invariant under any factorization, including forward, backward, or arbitrary permutations. This result establishes a rigorous theoretical foundation for studying how LLMs learn from data and defines principled protocols for empirical evaluation. Applying these protocols, we show that prior studies examining ordering effects suffer from critical methodological flaws. We retrain GPT-2 models across forward, backward, and arbitrary permuted orders on scientific text. We find systematic deviations from theoretical invariance across all orderings with arbitrary permutations strongly deviating from both forward and backward models, which largely (but not completely) agreed with one another. Deviations were traceable to differences in self-attention, reflecting positional and locality biases in processing. Our theoretical and empirical results provide novel avenues for understanding positional biases in LLMs and suggest methods for detecting when LLMs' probability distributions are inconsistent and therefore untrustworthy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。