arXiv:2507.22918cs.CLcs.LG2025-07被引 4

不同规模大模型在中间层形成相似语义特征,揭示了模型内部表征的通用性。

Semantic Convergence: Investigating Shared Representations Across Scaled LLMs

  • 用稀疏自编码器分析模型残差流激活,对齐单义特征并比较空间结构。
  • 中间层特征重叠度最高,四倍规模差异下仍保持显著语义一致性。
  • 结果支持跨模型可解释性基础,适合研究模型内部机制的学者参考。

我们研究了Gemma-2语言模型(Gemma-2-2B和Gemma-2-9B)中的特征通用性,探究规模相差四倍的模型是否仍收敛于相似的内部概念。通过稀疏自编码器(SAE)字典学习流程,对每个模型残差流激活进行SAE处理,利用激活相关性对齐所得单义特征,并通过SVCCA与RSA比较匹配后的特征空间。结果显示,中间层具有最强的重叠性,而早期和晚期层则显著不相似。初步实验将分析扩展至多标记子空间,表明语义相近的子空间对语言模型产生类似交互。这些结果强化了大语言模型尽管规模不同,仍能以广泛一致且可解释的特征刻画世界,支持通用性作为跨模型可解释性的基础。

原文摘要 · Abstract (English)

We investigate feature universality in Gemma-2 language models (Gemma-2-2B and Gemma-2-9B), asking whether models with a four-fold difference in scale still converge on comparable internal concepts. Using the Sparse Autoencoder (SAE) dictionary-learning pipeline, we utilize SAEs on each model's residual-stream activations, align the resulting monosemantic features via activation correlation, and compare the matched feature spaces with SVCCA and RSA. Middle layers yield the strongest overlap, while early and late layers show far less similarity. Preliminary experiments extend the analysis from single tokens to multi-token subspaces, showing that semantically similar subspaces interact similarly with language models. These results strengthen the case that large language models carve the world into broadly similar, interpretable features despite size differences, reinforcing universality as a foundation for cross-model interpretability.

模型可解释性特征通用性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。