发现语言模型用可解释的线性表示编码概念层级关系。
Linear Representations of Hierarchical Concepts in Language Models

- 为不同层级和领域训练特定线性变换,分析层次关系表征。
- 在单一领域内能线性恢复层级关系,跨领域迁移效果有限。
- 层次信息集中在低维、领域特异的子空间,但结构高度相似。
我们研究语言模型内部表征中层次关系(如日本 ⊂ 东亚 ⊂ 亚洲)的编码方式与程度。基于线性关系概念,我们为每个层次深度和语义领域训练特定线性变换,并通过比较这些变换来刻画层次关系相关的表征差异。相比以往对语言模型层次表征几何的研究,本工作涵盖多标记实体和跨层表示。在多个领域中学习这些变换,并评估域内泛化能力及跨域迁移表现。实验表明,在单一领域内,层次关系可从模型表征中线性恢复。进一步分析发现,层次信息编码于相对低维的子空间,且该子空间具有领域特异性。主要结果是:尽管子空间不同,但层次表征在这些领域特异性子空间中高度相似。总体而言,所有实验模型均以高度可解释的线性形式编码概念层级。
原文摘要 · Abstract (English)
We investigate how and to what extent hierarchical relations (e.g., Japan $\subset$ Eastern Asia $\subset$ Asia) are encoded in the internal representations of language models. Building on Linear Relational Concepts, we train linear transformations specific to each hierarchical depth and semantic domain, and characterize representational differences associated with hierarchical relations by comparing these transformations. Going beyond prior work on the representational geometry of hierarchies in LMs, our analysis covers multi-token entities and cross-layer representations. Across multiple domains we learn such transformations and evaluate in-domain generalization to unseen data and cross-domain transfer. Experiments show that, within a domain, hierarchical relations can be linearly recovered from model representations. We then analyze how hierarchical information is encoded in representation space. We find that it is encoded in a relatively low-dimensional subspace and that this subspace tends to be domain-specific. Our main result is that hierarchy representation is highly similar across these domain-specific subspaces. Overall, we find that all models considered in our experiments encode concept hierarchies in the form of highly interpretable linear representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。