发现大模型隐空间特征分布高度相似,可跨模型迁移解释技术。
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
- 用稀疏自编码器解耦模型特征,比较不同模型间特征空间的几何关系。
- 跨模型特征空间在旋转不变变换下相似度高达0.87以上,证明空间通用性。
- 适合研究模型可解释性迁移、跨模型干预与控制的科研人员。
大语言模型(LLMs)的通用性假设认为,不同模型在隐空间中会收敛到相似的概念表示。若该假设成立,将有助于实现机制可解释性技术在不同模型间的泛化。以往研究聚焦于模型是否学习到相同的特征(即激活特定概念的内部表示),但因神经元常具多义性,难以直接比较。为此,研究者采用稀疏自编码器(SAEs)将神经元分解为对应单一概念的独立特征。本文提出新的通用性假设——类比特征通用性:即使不同模型的SAE学到的特征表示不同,其特征空间在旋转不变变换下仍应相似。为验证此假设,我们通过激活相关性配对不同模型的SAE特征,并利用表征相似性度量分析配对特征间的空间关系。实验表明,多个大语言模型的SAE特征空间具有高度相似性,支持特征空间通用性的存在。
原文摘要 · Abstract (English)
The Universality Hypothesis in large language models (LLMs) claims that different models converge towards similar concept representations in their latent spaces. Providing evidence for this hypothesis would enable researchers to exploit universal properties, facilitating the generalization of mechanistic interpretability techniques across models. Previous works studied if LLMs learned the same features, which are internal representations that activate on specific concepts. Since comparing features across LLMs is challenging due to polysemanticity, in which LLM neurons often correspond to multiple unrelated features rather than to distinct concepts, sparse autoencoders (SAEs) have been employed to disentangle LLM neurons into SAE features corresponding to distinct concepts. In this paper, we introduce a new variation of the universality hypothesis called Analogous Feature Universality: we hypothesize that even if SAEs across different models learn different feature representations, the spaces spanned by SAE features are similar, such that one SAE space is similar to another SAE space under rotation-invariant transformations. Evidence for this hypothesis would imply that interpretability techniques related to latent spaces, such as steering vectors, may be transferred across models via certain transformations. To investigate this hypothesis, we first pair SAE features across different models via activation correlation, and then measure spatial relation similarities between paired features via representational similarity measures, which transform spaces into representations that reveal hidden relational similarities. Our experiments demonstrate high similarities for SAE feature spaces across various LLMs, providing evidence for feature space universality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。