用张量积探针发现语言模型中线性表示的深层结构
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions

- 用张量积探针解析线性方向间的关联结构
- 发现棋盘状态可分解为方格、颜色嵌入与绑定矩阵的组合
- 线性探针结果可直接从张量积参数还原,揭示其底层结构
尽管研究者已发现概念在语言模型中以线性方向表征,但单个线性方向无法捕捉关系结构。为理解这一矛盾,我们研究了一个具有已知线性表征但训练于高度结构化领域(奥赛罗棋)的模型。该模型内部棋盘状态表征可被线性解码,但我们发现了张量积表示(TPRs)的附加结构。通过训练TPR探针,我们恢复了线性探针间的共享结构,实现对正方形嵌入、颜色嵌入和绑定矩阵的因子分解,从而构建模型的棋盘状态表示。我们发现探针权重中存在几何特征,与棋盘结构一致;更重要的是,线性探针可直接从TPR探针参数中重构。这表明线性表示可能是更复杂底层表示的投影。
原文摘要 · Abstract (English)
While researchers are finding concepts represented as linear directions in language models, a bag of linear directions fails to capture relational structure. To better understand this dichotomy, we study a model with known linear representations, but trained in a highly structured domain -- the board game Othello. While the model's internal board-state representation is linearly decodable, we find additional structure in the form of tensor product representations (TPRs). We train TPR probes to recover shared structure amongst the linear probes, yielding a factorization into square-embeddings, color-embeddings, and a binding matrix that composes them to construct the model's board-state representation. We find geometric signatures within the weights of our TPR probe that align with the structure of the board, but perhaps more importantly, that the linear probes can be recovered directly from the parameters of our TPR probe. Our findings suggest that directional representations may be projections of more structured underlying representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。