arXiv:2608.15448cs.CLcs.LG2026-08

语言模型在模糊的下一个词分布上学习困难,影响生成质量。

Language models suffer from a curse of ambiguity

论文配图:Language models suffer from a curse of ambiguity
图 1 · 摘自论文原文
  • 发现模糊性诅咒:越模糊的词分布越难准确学习。
  • 模糊分布需更大容量、更多训练步数,且易受采样噪声干扰。
  • 适用于评估模型输出可信度,尤其对高风险应用有指导意义。

大型语言模型越来越依赖采样来驱动自我改进,因此其学习到的概率分布保真度至关重要。然而,并非所有分布都同样易于学习。本文揭示了‘模糊性诅咒’:在大语言模型及所有生成离散概率分布的神经网络中,下一个词分布越模糊,学习就越困难。通过系统的理论分析,我们发现这一现象源于架构和学习机制的根源——模糊分布需要更大的模型容量存储、更长的嵌入表示、更多训练步骤拟合,并放大采样噪声。我们在具有可控真实标签的合成任务上验证了这些发现,并在真实数据训练的语言模型中观察到相同特征。结果为理解大语言模型的统计能力提供了新视角,并提供了一个实用框架,用于判断何时可信任其输出分布。

原文摘要 · Abstract (English)

Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever. Yet, not all distributions are equally easy to learn. In this work, we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately. Through an extensive theoretical analysis, we trace this curse to architectural and learning roots. More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise. We validate these findings on synthetic tasks with controlled ground truth and observe the same signatures in language models trained on real data. Our results provide a new perspective on the statistical capabilities of large language models and a practical framework for when to trust their output distribution.

语言模型模糊性分布学习可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。