arXiv:2511.12768cs.CLcs.AI2025-11

小模型训练中也存在突变式能力跃迁,早于损失下降显现。

Evidence of Phase Transitions in Small Transformer-Based Language Models

  • 用词汇使用和统计规律探测训练过程中的突变点。
  • 在训练早期即出现能力跃迁,与损失曲线无关。
  • 适合研究模型动态和非线性训练机制的学者。

相变现象被视作大语言模型涌现能力的起源,即当模型规模突破临界阈值时,新能力会突然出现。此前工作(如 Wei 等)在模型与数据缩放下通过计算量对数变换揭示了此类现象。本文提出三个互补问题:(1) 相变是否仅存在于大模型?(2) 能否在原始线性训练空间中直接检测到?(3) 是否能在训练早期出现?为此,我们在字符级语料上训练一个小的 GPT 风格变压器模型,分析词汇使用演化过程,追踪平均词长、正确/错误词数及词汇多样性变化。结合泊松与超泊松统计,量化词汇连接与重组方式。结果揭示了一个明显的过渡点。值得注意的是,该现象未出现在标准损失或验证曲线中,但通过词汇与统计探针可清晰识别。研究发现,相变重组是语言模型训练的普遍特征,即便在中小模型中亦可直接在原始线性空间中观测,且发生于连贯性刚出现的早期阶段。这一视角为理解语言模型训练的非线性动态提供了新见解,并强调了定制化度量的重要性。

原文摘要 · Abstract (English)

Phase transitions have been proposed as the origin of emergent abilities in large language models (LLMs), where new capabilities appear abruptly once models surpass critical thresholds of scale. Prior work, such as that of Wei et al., demonstrated these phenomena under model and data scaling, with transitions revealed after applying a log scale to training compute. In this work, we ask three complementary questions: (1) Are phase transitions unique to large models, or can they also be observed in small transformer-based language models? (2) Can such transitions be detected directly in linear training space, rather than only after log rescaling? and (3) Can these transitions emerge at early stages of training? To investigate, we train a small GPT-style transformer on a character-level corpus and analyze the evolution of vocabulary usage throughout training. We track the average word length, the number of correct versus incorrect words, and shifts in vocabulary diversity. Building on these measures, we apply Poisson and sub-Poisson statistics to quantify how words connect and reorganize. This combined analysis reveals a distinct transition point during training. Notably, these transitions are not apparent in standard loss or validation curves, but become visible through our vocabulary- and statistics-based probes. Our findings suggest that phase-transition reorganizations are a general feature of language model training, observable even in modest models, detectable directly in linear training space, and occurring surprisingly early as coherence emerges. This perspective provides new insight into the nonlinear dynamics of language model training and underscores the importance of tailored metrics for uncovering phase transition behaviors

语言模型相变训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。