解释语言模型中线性规律为何普遍存在,发现要么所有等效模型都有,要么全都没有。
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
- 从分布等价性出发,证明线性性质在模型间可识别
- 揭示特定线性关系在等效模型中或全有或全无的特性
- 适合研究模型内部表征与数学规律的学者参考
我们探讨可识别性是否能解释语言模型中广泛存在的线性特性,例如'easy'与'easiest'的表征差向量,与'lucky'和'luckiest'的差向量平行。核心问题是:若一个模型具备某种线性性质,是否所有生成相同分布的模型也必然具备?为此,我们首先证明了一个可识别性结果,推广了先前研究对多样性假设的依赖;其次,基于对关系线性性的改进定义(Paccanaro and Hinton, 2001;Hernandez et al., 2024),展示多种线性概念可纳入分析框架;最后,在适当条件下,证明这些线性性质在分布等价的下一词预测器中要么全部存在,要么全不存在。
原文摘要 · Abstract (English)
We analyze identifiability as a possible explanation for the ubiquity of linear properties across language models, such as the vector difference between the representations of "easy" and "easiest" being parallel to that between "lucky" and "luckiest". For this, we ask whether finding a linear property in one model implies that any model that induces the same distribution has that property, too. To answer that, we first prove an identifiability result to characterize distribution-equivalent next-token predictors, lifting a diversity requirement of previous results. Second, based on a refinement of relational linearity [Paccanaro and Hinton, 2001; Hernandez et al., 2024], we show how many notions of linearity are amenable to our analysis. Finally, we show that under suitable conditions, these linear properties either hold in all or none distribution-equivalent next-token predictors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。