arXiv:2501.01638cs.CLcs.AI2025-01被引 1

用非遍历理论解释大模型能力为何随规模突然涌现。

A non-ergodic framework for understanding emergent capabilities in Large Language Models

  • 基于相邻可能理论构建非遍历框架,揭示能力涌现机制。
  • 实验证明能力通过语义空间相变离散出现,受约束交互驱动。
  • 为模型架构设计提供理论指导,适合研究模型涌现的学者。

大语言模型在规模增大时会出现意外涌现能力,但缺乏理论解释。本文证明语言模型本质上是非遍历系统,并基于斯图尔特·考夫曼的相邻可能理论(TAP)构建数学框架,解释能力如何涌现。资源受限的TAP方程揭示了架构、训练和上下文约束如何通过语义空间中的相变共同塑造模型能力。通过对三种不同语言模型的实验验证,发现能力通过离散相变涌现,受约束交互与路径依赖探索引导。该框架为理解语言模型中的涌现现象提供了理论基础,并指导可引导能力涌现的架构设计。

原文摘要 · Abstract (English)

Large language models have emergent capabilities that come unexpectedly at scale, but we need a theoretical framework to explain why and how they emerge. We prove that language models are actually non-ergodic systems while providing a mathematical framework based on Stuart Kauffman's theory of the adjacent possible (TAP) to explain capability emergence. Our resource-constrained TAP equation demonstrates how architectural, training, and contextual constraints interact to shape model capabilities through phase transitions in semantic space. We prove through experiments with three different language models that capacities emerge through discrete transitions guided by constraint interactions and path-dependent exploration. This framework provides a theoretical basis for understanding emergence in language models and guides the development of architectures that can guide capability emergence.

大模型涌现能力非遍历理论框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。