arXiv:2608.29034cs.CLcs.AI2026-08

用张量积表征统一解释大模型的多种可解释性方法。

A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

论文配图:A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability
图 1 · 摘自论文原文
  • 以张量积表征作为统一框架,描述语言模型中成分结构的向量表示。
  • 数学与实证验证:四种主流可解释性方法均可由张量积推导得出。
  • 适用于从小模型到大模型的广泛场景,为模型理解提供统一视角。

众多语言模型可解释性方法已揭示其内部运作机制,但这些方法及其结论相互孤立。本文提出使用张量积表征(TPRs)作为统一假设,认为组合结构可通过填充-角色绑定在向量空间中实现。我们从数学和实证两方面证明,加法类比、线性探测、稀疏自编码器和激活修补等方法均可由TPRs推导。数学上,这些方法均是TPR的特例;实证上,在从小型玩具模型到大语言模型的多种架构中,基于TPR构造的变体性能与标准方法相当。本工作标志着迈向可解释性理想目标的重要一步:不仅解释个体现象,更阐明不同方法间的内在关联。

原文摘要 · Abstract (English)

A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language models be, such that they give rise to all our interpretations? In this work, we propose using Tensor Product Representations (TPRs) as a unifying hypothesis. TPRs give a concrete proposal for how compositional structure could be represented in vector space --- as filler-role bindings. We show, both mathematically and empirically, that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching. Mathematically, we show that these methods can all be derived from TPRs. Empirically, we apply the derivations to a range of different models --- from small toy models to LLMs --- to construct instances of each of the above interpretability methods; these constructed variants perform comparably to their standard variants. We view this work as a step toward what interpretability will ideally provide: a unified account of the nature of neural networks, corroborated not just by individual observations but also by an explanation of the connections between them.

可解释性张量积语言模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。