不同大模型在相似层位的激活值高度相似,揭示了深层语义结构的共性。
Layers at Similar Depths Generate Similar Activations Across LLM Architectures
- 对比24个独立训练的大模型,分析各层激活值的最近邻关系。
- 相同层位的模型间激活模式高度一致,跨模型共享几何结构。
- 适合关注模型内部表征通用性与可解释性的研究者阅读。
我们研究了24个开源大语言模型在不同层位上的激活值所诱导的最近邻关系。发现:1)同一模型内,不同层的激活近邻关系随层变化;2)不同模型中对应层的激活近邻关系近似共享。第二点表明这些关系并非随机,而是跨模型存在的共性;第一点则说明不存在全模型通用的单一近邻结构。二者共同提示:大模型在从浅层到深层的演进过程中,生成了一套逐渐变化的激活几何结构,但整个演化过程在不同架构间基本保持一致,仅通过拉伸或压缩适配具体结构。
原文摘要 · Abstract (English)
How do the latent spaces used by independently-trained LLMs relate to one another? We study the nearest neighbor relationships induced by activations at different layers of 24 open-weight LLMs, and find that they 1) tend to vary from layer to layer within a model, and 2) are approximately shared between corresponding layers of different models. Claim 2 shows that these nearest neighbor relationships are not arbitrary, as they are shared across models, but Claim 1 shows that they are not "obvious" either, as there is no single set of nearest neighbor relationships that is universally shared. Together, these suggest that LLMs generate a progression of activation geometries from layer to layer, but that this entire progression is largely shared between models, stretched and squeezed to fit into different architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。