arXiv:2511.12869cs.LGcs.AI2025-11被引 21

揭示大模型规模扩展的五大理论极限,解释为何再扩大也难突破瓶颈。

On the Fundamental Limits of LLMs at Scale

  • 从计算、信息、学习三方面建立理论框架,证明规模扩张存在不可逾越的天花板。
  • 指出模型在长文本压缩、推理退化、事实检索上存在根本性误差,无法通过单纯扩容解决。
  • 适合对模型局限性感兴趣的研究者和工程团队,为优化提供理论依据。

大规模语言模型(LLMs)虽受益于规模增长,但其提升受限于五个根本性问题:幻觉、上下文压缩、推理退化、检索脆弱性和多模态错位。现有综述仅描述现象,缺乏将这些现象与计算、信息和学习的基础极限相联系的严谨理论。本文提出一个统一且基于证明的框架,形式化了大模型规模扩展的内在理论上限。首先,可计算性与不可计算性表明误差不可避免:对于任何可枚举的模型族,对角化论证保证存在某些输入会使模型必然失败;不可判定任务(如停机类问题)会导致所有可计算预测器产生无限错误集。其次,信息论与统计约束限制了可实现的准确率,有限描述长度引发压缩误差,长尾事实知识需要难以企及的样本复杂度。第三,几何与计算效应导致长上下文被压缩至远低于名义尺寸,原因包括位置训练不足、编码衰减和softmax拥挤。我们进一步证明,基于似然的训练偏好模式补全而非推理,检索在词元限制下受语义漂移和耦合噪声影响,多模态扩展继承浅层跨模态对齐。各部分结合定理与实证证据,阐明规模扩展何时有效、何时饱和、何时无法推进,给出理论基础与实践缓解路径,如受限预言机检索、位置课程学习、稀疏或分层注意力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have benefited enormously from scaling, yet these gains are bounded by five fundamental limitations: (1) hallucination, (2) context compression, (3) reasoning degradation, (4) retrieval fragility, and (5) multimodal misalignment. While existing surveys describe these phenomena empirically, they lack a rigorous theoretical synthesis connecting them to the foundational limits of computation, information, and learning. This work closes that gap by presenting a unified, proof-informed framework that formalizes the innate theoretical ceilings of LLM scaling. First, computability and uncomputability imply an irreducible residue of error: for any computably enumerable model family, diagonalization guarantees inputs on which some model must fail, and undecidable queries (e.g., halting-style tasks) induce infinite failure sets for all computable predictors. Second, information-theoretic and statistical constraints bound attainable accuracy even on decidable tasks, finite description length enforces compression error, and long-tail factual knowledge requires prohibitive sample complexity. Third, geometric and computational effects compress long contexts far below their nominal size due to positional under-training, encoding attenuation, and softmax crowding. We further show how likelihood-based training favors pattern completion over inference, how retrieval under token limits suffers from semantic drift and coupling noise, and how multimodal scaling inherits shallow cross-modal alignment. Across sections, we pair theorems and empirical evidence to outline where scaling helps, where it saturates, and where it cannot progress, providing both theoretical foundations and practical mitigation paths like bounded-oracle retrieval, positional curricula, and sparse or hierarchical attention.

大模型理论极限推理退化信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。