arXiv:2606.28105cond-mat.dis-nncond-mat.stat-mech2026-06被引 1

揭示随机语言模型在大尺度下的相变规律,解释语言统计的长程依赖。

Scaling limit of the Random Language Model

论文配图:Scaling limit of the Random Language Model
图 1 · 摘自论文原文
  • 基于大偏差原理与半退火近似,建立语法生成的数学框架。
  • 发现临界点x=1/8处规则使用发生凝聚,熵在x=1/2时开始下降。
  • 为自然语言统计和大模型行为提供统一理论基础,适合理论计算语言学研究者。

我们发展了随机语言模型(RLM)的定量理论,该模型是一类随机上下文无关文法的集合。在隐藏符号数N趋于无穷,而文法温度ε̃_d趋于零、固定x = ε̃_d log N的极限下,模型可通过规则使用模式的大偏差原理进行控制描述。半退火近似将问题映射为具有非平凡组合结构的随机能量模型。我们证明,当x_c = 1/8时,模型出现凝聚相变:低于此值时规则使用集中,语言统计表现出对语料长度的非平凡依赖。在x = 1/2处出现第二个特征尺度,标志着熵从最大值开始减少。在这些区域间,我们推导出不同规则数量、熵及相关可观测量的显式标度律,揭示了由语法规模、语料长度与温度共同调控的区分标度、饱和与临界行为。该理论解决了此前关于热力学相变存在的模糊性,并解释了大N极限缓慢收敛源于log N依赖。它还为自然语言统计和大语言模型行为中普遍统计性质的涌现,提供了统一生成框架。

原文摘要 · Abstract (English)

We develop a quantitative theory of the Random Language Model (RLM), an ensemble of stochastic context-free grammars, in a scaling limit where the number of hidden symbols $N \to \infty$ while the grammar temperature $\tildeε_d \to 0$ at fixed $x = {\tildeε}_d \log N$. In this limit, the model admits a controlled description based on a large-deviation principle over rule-usage patterns. A semi-annealed approximation maps the problem to a class of Random Energy Models with nontrivial combinatorics. We show that the RLM exhibits a condensation transition at a critical value $x_c=1/8$, below which rule usage concentrates and language statistics acquire a nontrivial dependence on corpus length. A second characteristic scale at $x=1/2$ marks the onset of entropy reduction from its maximal value. Across these regimes, we derive explicit scaling laws for the number of distinct rules, entropy, and related observables, identifying distinct scaling, saturation, and critical regimes controlled by the interplay of grammar size, corpus length, and temperature. The theory resolves previous ambiguities regarding the existence of a thermodynamic transition and explains the slow approach to the large-$N$ limit as a consequence of the dependence on $\log N$. It further provides a unified framework in which universal statistical properties of language emerge from typical realizations of generative grammars, with implications for both natural language statistics and the behavior of large language models.

语言模型相变统计物理生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。