arXiv:2508.01916cs.LGcs.AI2025-08被引 6

无监督学习发现神经网络中可解释的子空间

Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning

  • 通过邻居距离最小化方法,无监督地挖掘出非基对齐的可解释子空间
  • 在GPT-2中发现子空间与已知电路变量高度相关,验证其语义一致性
  • 适用于大规模模型,适合研究模型内部机制与可解释性

理解神经模型内部表示是机制可解释性的核心问题。由于表示空间维度高,可能编码输入的多个方面。这些不同方面是否以独立子空间形式组织?能否完全无监督地发现这些“自然”子空间?令人惊讶的是,我们确实可以实现这一目标,并通过看似无关的训练目标发现可解释子空间。我们的方法——邻居距离最小化(NDM)——以无监督方式学习非基对齐子空间。定性分析显示,这些子空间在许多情况下具有可解释性,且所编码的信息在不同输入间共享相同抽象概念,类似于模型使用的“变量”。我们在GPT-2中进行定量实验,结果表明子空间与电路变量存在强关联。此外,我们还提供了扩展至20亿参数模型的证据,成功分离出负责上下文与参数知识路由的独立子空间。总体而言,本研究为理解模型内部结构和构建可解释电路提供了新视角。

原文摘要 · Abstract (English)

Understanding internal representations of neural models is a core interest of mechanistic interpretability. Due to its large dimensionality, the representation space can encode various aspects about inputs. To what extent are different aspects organized and encoded in separate subspaces? Is it possible to find these ``natural'' subspaces in a purely unsupervised way? Somewhat surprisingly, we can indeed achieve this and find interpretable subspaces by a seemingly unrelated training objective. Our method, neighbor distance minimization (NDM), learns non-basis-aligned subspaces in an unsupervised manner. Qualitative analysis shows subspaces are interpretable in many cases, and encoded information in obtained subspaces tends to share the same abstract concept across different inputs, making such subspaces similar to ``variables'' used by the model. We also conduct quantitative experiments using known circuits in GPT-2; results show a strong connection between subspaces and circuit variables. We also provide evidence showing scalability to 2B models by finding separate subspaces mediating context and parametric knowledge routing. Viewed more broadly, our findings offer a new perspective on understanding model internals and building circuits.

可解释性子空间分解无监督学习模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。