语言模型中的概念层次结构源于词语共现统计的谱特性。
Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence

- 基于词语共现频率构建词向量矩阵,通过谱分析揭示层次结构。
- 主特征向量按粗到细顺序分离语义分支,与词典树结构一致。
- 该现象在word2vec和Gemma 2B中均存在,说明其来自统计本质。
我们提出一种分布理论,解释超类关系(即“是”关系)如何在语言表征中以几何方式编码。基于词语在WordNet超类图上距离越近则共现越频繁的实证假设,我们理论刻画了word2vec词向量嵌入的共现矩阵谱分布。在共现核满足温和正性与衰减条件下,证明主特征向量先分离广义分类枝干,再逐步细化子分支,形成从粗到细的分层分裂几何结构,与树状结构高度吻合。我们在多个采样的WordNet子树中验证了这一预测,并发现该谱签名在Gemma 2B未嵌入中也显著成立。结果表明,大型语言模型中的层次概念几何无需专门的层级机制,而是由成对词语统计的谱结构自然涌现。
原文摘要 · Abstract (English)
We propose a distributional theory of how hypernymy -- the ``is-a'' relation between general and specific concepts -- is encoded geometrically in language representations. Starting from the empirically verified assumption that words closer on the WordNet hypernym graph co-occur more often, we characterize theoretically the spectrum of the resulting embedding Gram matrix of word2vec embeddings. Under mild positivity and decay conditions on the co-occurrence kernel, we prove that the leading eigenvectors first separate broad taxonomic branches and then progressively finer sub-branches, producing a \emph{hierarchical splitting geometry} with a coarse-to-fine spectral organization that mirrors the tree. We confirm these predictions in word2vec embeddings across many sampled WordNet subtrees, and show that the same signature extends strikingly well to Gemma 2B unembeddings. Our results indicate that hierarchical concept geometry in LLMs need not reflect a hierarchy-specific functional mechanism, but emerges from the spectral structure of pairwise word statistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。