arXiv:2602.15997cs.LGcs.AI2026-02被引 1

模型能力提升前,先出现隐藏状态坍缩,再恢复,这一几何变化可预测难任务的性能突破。

The Geometric Anatomy of Capability Acquisition in Transformers

  • 追踪模型各层几何变化与线性探测器,发现能力提升前隐藏状态先坍缩后恢复。
  • 在2.8B参数模型上,逻辑推理任务提前约4.9万步出现几何变化,易任务则无明显前置信号。
  • 只有秩(rank)能可靠预示困难任务的能力获取,且该现象在大模型中依然存在。

神经网络在训练中获得新能力,但能力涌现前的内部变化机制尚不清晰。本文系统追踪了6种Transformer规模(405K–151M参数)、8类算法任务(共144个任务×层级×模型组合)以及3个Pythia语言模型(160M–2.8B参数)中的几何特征与线性探测结果。所有设置下,模型表示均先经历低维坍缩,随后恢复,能力提升才发生。线性探测显示,隐藏状态在行为表现提升前已包含任务相关信息。坍缩程度具有任务特异性,且自顶向下传播;在所测几何指标中,仅秩(rank)能在困难任务中可靠地预示能力获取。是否可观测到该前置信号取决于任务难度与模型容量的相对关系:对于困难任务,几何变化先于行为变化(如2.8B模型上的逻辑推理任务,差距约49K训练步);而对于简单任务,两者几乎同步,无显著前置信号。这表明小模型中观察到的几何模式可在大模型中延续,前提是任务仍对模型构成挑战。

原文摘要 · Abstract (English)

Neural networks gain capabilities during training, but the internal changes that precede capability acquisition are not well understood. In particular, the relationship between geometric change and behavioral change, and the effect of task difficulty and model scale on that relationship, is unclear. We track geometric measures and linear probes across six transformer sizes (405K--151M parameters), eight algorithmic tasks (144 task$\times$level$\times$model combinations), and three Pythia language models (160M--2.8B). Across all settings, representations first collapse to a low-dimensional state, then recover, and only then does behavioral performance improve. Linear probes show that the model's hidden states already contain task-relevant information before the model can act on it. The collapse floor is task-specific, the collapse propagates top-down through the network, and of the geometric measures tested, only \rankme reliably precedes capability acquisition for hard tasks. Whether this precursor is detectable depends on task difficulty relative to model capacity. For hard tasks, there is a clear gap: geometry changes first, behavior follows. For easy tasks, the model learns so quickly that both happen simultaneously and no precursor is detectable. On Pythia-2.8B, a logical deduction task that is genuinely hard for the model shows a precursor gap of ${\sim}$49K training steps, while easy benchmarks show none. This suggests that geometric patterns observed in small proxy models can persist at larger scale when the task remains difficult relative to model capacity.

Transformer能力涌现几何分析模型规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。