语言模型可靠性有理论上限,无法仅靠扩大规模突破。
Information-Theoretic Limits of Reliability and Scaling in Language Models

- 从信息论出发,指出生成任务的可靠性受上下文可解不确定性限制。
- 模型性能受限于训练数据或模型容量中更稀缺的资源,且存在不可逾越的上限。
- 解释了检索增强、灾难性遗忘等现象,适合研究模型极限的学者。
大型语言模型(LLMs)常被评估为只要规模足够大,任何任务都能达到完美可靠性。我们证明这一假设在信息论上不成立。每个生成任务都有一个可靠性天花板,由可观测上下文能消除的输出不确定性决定。该差距可分解为可通过更多上下文关闭的可解析部分和任务固有的主观模糊部分。自回归生成进一步以依赖核速率降低此天花板,依赖核量化输出中词元间的相关性。基于这两个基本要素,我们推导出首个原理性缩放定律:LLM性能由更稀缺的资源——训练数据或模型容量——所瓶颈。该定律将Chinchilla缩放定律作为特例,并提供了何时缩放能提升可靠性的结构性解释。超越缩放,我们的框架统一了多种实际现象,如检索增强的优势及灾难性遗忘的谱机制。本工作形式化了跨领域模型性能的资源-复杂度权衡,为生成式语言模型的性能极限提供了统一理论。
原文摘要 · Abstract (English)
Large language models (LLMs) are evaluated as though perfect reliability is achievable for any task given sufficient scale. We show this assumption is information-theoretically unjustified. Every generative task has a reliability ceiling that no model can exceed, determined by how much output uncertainty is resolvable from observable context. The gap decomposes into a resolvable component closable with additional context and a subjective component inherent to task ambiguity. Autoregressive generation further degrades this ceiling at a rate governed by the task's dependency kernel, which quantifies inter-token correlations in the output. From these two primitives, we derive a first-principles scaling law where LLM performance is bottlenecked by the scarcer resource: training data or model capacity. This law recovers the Chinchilla scaling law as a special case and provides a structural account of when scaling improves reliability. Beyond scaling, our framework unifies diverse practical phenomena, such as the benefits of retrieval-augmentation and the spectral mechanics of catastrophic forgetting. Our work formalizes the resource-complexity tradeoffs that govern model performance across domains, offering a unified theory of performance limits in generative language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。