语言模型在单标签训练中仍能发现隐含语义结构,因早期训练中出现临时的语义几何。
Structure Before Collapse: Transient semantic geometry in next-token prediction

- 通过合成数据验证:即使输入有隐含语义特征,模型仍会自发聚类。
- 语义结构在训练初期显现,但随容量和时间增长逐渐消失于对称状态。
- 适合研究模型内部表征演化与梯度下降机制的人阅读。
神经坍缩理论预测,平衡的一热分类会使模型表示等距分布,仅依赖输出标签而忽略输入语义相似性。这引发疑问:尽管下一个词预测模型主要以一热标签训练(上下文极少重复且标签不同),为何仍能学习到潜在的结构特征?例如,在'玛丽打破了___'中,模型能识别出下个词可能为中等尺寸、坚硬、无生命的名词。当共现统计趋于一热稀疏、不同上下文间无共享目标词时,梯度下降如何发现这种类别语义结构?我们设计了三个受控合成场景,其中输入包含隐含语义因子但被映射到不同一热标签。结果表明,语义几何在训练初期即出现,表示按共同属性聚类,虽无显式监督。该结构是瞬态的:在足够容量与时间下,模型最终趋向预测的对称状态,所有表示等距分离。我们通过格拉姆矩阵分析研究这一相变,并提出对常用无约束特征模型的初步改进以捕捉此涌现语义几何。
原文摘要 · Abstract (English)
Neural Collapse predicts that balanced one-hot classification pushes model representations to be equally far from each other; a symmetric configuration that depends only on the output label and ignores any semantic similarity in the inputs. This creates a puzzle: next-token prediction language models are trained predominantly (as context length increases) with one-hot labels: the same context is very unlikely to appear twice in training with different labels. However, they clearly learn latent structural features. That is, despite the one-hot training regime, a language model's contextual embeddings represent the fact that the next word in ''Mary broke the ___'' is likely to be filled by tokens in the latent classes of a) medium-sized, b) rigid, c) inanimate nouns. How does gradient descent find such categorical semantic structure when co-occurrence statistics collapse to one-hot sparsity, eliminating any shared next-tokens among different contexts? To investigate this tension we identify three synthetic controlled settings where inputs have latent semantic factors but are mapped to distinct one-hot labels. We find that semantic geometry emerges early in training, and that representations cluster by shared attributes despite receiving no explicit supervision to do so. This structure is transient: with sufficient capacity and time, the model eventually reaches the predicted symmetric state where all representations are equally separated. We study this phase transition through Gram matrix analysis and propose a preliminary modification to the commonly used unconstrained features model to capture the emergent semantic geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。