用几何维度揭示语言模型如何学习组合性原理
Geometric Signatures of Compositionality Across a Language Model's Lifetime
- 通过内在维度衡量语言模型表征的复杂度
- 数据组合性越强,表征内在维度越低,反映语言简洁性
- 非线性维度捕捉语义组合,线性维度反映表面模式
由于语言的组合性,少量句法规则和有限词汇即可生成无限句子。这意味着语言虽看似高维,实则可由较少自由度解释。当前语言模型是否反映了这种由组合性带来的内在简洁性,仍是个开放问题。我们从几何视角出发,将数据集的组合性程度与语言模型表征的内在维度(ID)关联起来,该指标衡量特征复杂度。结果表明:数据集的组合性确实反映在表征的内在维度中;且这一关系源于训练过程中学习到的语言特征。此外,分析揭示了非线性与线性维度之间的显著差异:前者编码语义层面的组合性,后者反映表面形式的组合规律。
原文摘要 · Abstract (English)
By virtue of linguistic compositionality, few syntactic rules and a finite lexicon can generate an unbounded number of sentences. That is, language, though seemingly high-dimensional, can be explained using relatively few degrees of freedom. An open question is whether contemporary language models (LMs) reflect the intrinsic simplicity of language that is enabled by compositionality. We take a geometric view of this problem by relating the degree of compositionality in a dataset to the intrinsic dimension (ID) of its representations under an LM, a measure of feature complexity. We find not only that the degree of dataset compositionality is reflected in representations' ID, but that the relationship between compositionality and geometric complexity arises due to learned linguistic features over training. Finally, our analyses reveal a striking contrast between nonlinear and linear dimensionality, showing they respectively encode semantic and superficial aspects of linguistic composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。