揭示语言模型如何通过优化自动组织语义结构
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
- 用SVD隐式分解上下文-词共现矩阵,学习语义因子
- 大奇异值对应概念先学,形成从泛到精的语义层次
- 基于符号模式聚类,可识别语法、实体等语义类别
我们研究了下一词预测(NTP)优化如何使语言模型从文本中提取并组织语义结构。基于可处理的数学模型与受控合成数据,分析发现:尽管模型未显式构建,但其学习的词嵌入和上下文嵌入会收敛到一个中心化共现矩阵的SVD因子,奇异向量通过符号模式编码潜在语义概念。我们证明,对应较大奇异值的概念在训练初期即被学习,形成自然的语义层级,宽泛类别先于细粒度类别出现。该发现启发了基于象限的聚类方法,通过组合概念符号以识别可解释的语义类别。我们在合成数据集和预训练语言模型上验证了结果,成功恢复了语法类别、命名实体类型及主题差异(如医疗、娱乐)。本工作连接了经典分布语义与神经坍缩几何,揭示了梯度优化如何隐式决定编码语义结构的矩阵表示及其分解方式。
原文摘要 · Abstract (English)
We investigate how next-token prediction (NTP) optimization leads language models to extract and organize semantic structure from text. Our analysis, based on a tractable mathematical model and controlled synthetic data, reveals that NTP implicitly guides models to factor a centered support matrix encoding context-to-next-token co-occurrence patterns via singular value decomposition (SVD). While models never explicitly construct this matrix, learned word and context embeddings converge to its SVD factors, with singular vectors encoding latent semantic concepts through their sign patterns. We demonstrate that concepts corresponding to larger singular values are learned earlier during training, yielding a natural semantic hierarchy where broad categories emerge before fine-grained ones. This insight motivates orthant-based clustering, a method that combines concept signs to identify interpretable semantic categories. We validate our findings on synthetic datasets and pretrained language models, recovering diverse semantic structures such as grammatical categories, named entity types, and topical distinctions (medical, entertainment). Our work bridges classical distributional semantics and neural collapse geometry, characterizing how gradient-based optimization implicitly determines both the matrix representation and factorization method that encode semantic structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。