arXiv:2503.04725cs.CLcs.AI2025-03NeurIPS被引 9

提出长上下文建模的互信息定律,揭示模型记忆需随文本变长而指数增长。

L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling

  • 基于双分组互信息定律,刻画长文本中多标记间独立依赖关系。
  • 发现模型历史状态需随上下文长度指数级增长才能有效建模。
  • 适用于Transformer与状态空间模型,可指导高效架构设计。

我们提出一个基于双分组互信息缩放律的通用理论框架,用于理解长上下文语言建模,并在自然语言中严格验证了该定律。结果表明,双分组互信息能捕捉不同于传统两点互信息的多标记交互,且其缩放行为独立于后者,为准确建模长序列提供了更完整的依赖表征。基于此缩放律,我们推导出长上下文语言建模(L²M)条件,对模型历史状态(负责存储过往信息的隐变量)的必要缩放给出了下界。我们在Transformer和状态空间模型上验证了该框架及其预测。本工作为理解长上下文建模提供了原则性基础,并可指导设计具备更强长上下文能力的高效架构,潜在应用不限于自然语言。

原文摘要 · Abstract (English)

We present a universal theoretical framework for understanding long-context language modeling based on a bipartite mutual information scaling law that we rigorously verify in natural language. We demonstrate that bipartite mutual information captures multi-token interactions distinct from and scaling independently of conventional two-point mutual information, and show that this provides a more complete characterization of the dependencies needed for accurately modeling long sequences. Leveraging this scaling law, we formulate the Long-context Language Modeling (L$^2$M) condition, which lower bounds the necessary scaling of a model's history state -- the latent variables responsible for storing past information -- for effective long-context modeling. We validate the framework and its predictions on transformer and state-space models. Our work provides a principled foundation to understand long-context modeling and to design more efficient architectures with stronger long-context capabilities, with potential applications beyond natural language.

语言建模互信息长序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。