arXiv:2502.01481cs.LGcs.CL2025-02被引 17

用内在熵解释大模型上下文长度的影响机制

Intrinsic Entropy of Context Length Scaling in LLMs

  • 提出用内在熵分析上下文长度对语言建模的影响
  • 发现训练数据量决定最优上下文长度
  • 为长上下文模型设计提供理论依据

近年来,长上下文语言模型受到广泛关注。已有研究探讨长上下文对语言模型性能的影响:部分研究指出无关长上下文会损害性能,而另一些则通过实验总结出相关长上下文带来的损失下降符合缩放定律。这需要更深入理解上下文长度如何影响语言建模。本文(1)提出使用「内在熵」来解释上下文长度对语言建模的影响;(2)在自然语言和合成数据上进行实验,验证了所提出的理论假设与推论。我们的理论框架可提供实用洞见,例如:训练数据集大小决定了最优上下文长度,并对某些情况下的上下文长度缩放进行边界约束。我们希望本工作能启发新的长上下文语言模型,以及未来关于语言模型物理特性的研究。

原文摘要 · Abstract (English)

Long Context Language Models have drawn great attention in the past few years. There has been work discussing the impact of long context on Language Model performance: some find that long irrelevant context could harm performance, while some experimentally summarize loss reduction by relevant long context as Scaling Laws. This calls for a more thorough understanding of how long context impacts Language Modeling. In this work, we (1) propose to use `Intrinsic Entropy' for explaining the impact of context length on language modeling; and (2) conduct experiments on natural language and synthetic data, validating our proposed theoretical assumptions and deductions. Our theoretical framework can provide practical insights such as establishing that training dataset size dictates an optimal context length and bounds context length scaling for certain cases. We hope our work may inspire new long context Language Models, as well as future work studying the physics of Language Models.

大模型上下文长度理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。