arXiv:2411.08166cs.LG2024-11被引 1

用神经元嵌入解析模型中的多义性,提升可解释性。

Tackling Polysemanticity with Neuron Embeddings

  • 通过内部表征与权重计算神经元嵌入,捕捉其语义行为
  • 在GPT2-small上验证,可量化神经元的多义程度
  • 适用于自动或人工分析,对稀疏自编码器评估有帮助

我们提出神经元嵌入,一种可识别神经元在典型数据样本中不同语义行为的表征,从而大幅简化下游的人工或自动解释。该方法应用于GPT2-small,并提供了可视化界面以探索结果。神经元嵌入基于模型内部表示和权重计算,具有领域与架构无关性,避免引入可能不反映真实计算过程的外部结构。我们阐述了如何利用神经元嵌入衡量神经元多义性,可用于更有效地评估稀疏自编码器(Sparse Auto-Encoders, SAEs)的性能。

原文摘要 · Abstract (English)

We present neuron embeddings, a representation that can be used to tackle polysemanticity by identifying the distinct semantic behaviours in a neuron's characteristic dataset examples, making downstream manual or automatic interpretation much easier. We apply our method to GPT2-small, and provide a UI for exploring the results. Neuron embeddings are computed using a model's internal representations and weights, making them domain and architecture agnostic and removing the risk of introducing external structure which may not reflect a model's actual computation. We describe how neuron embeddings can be used to measure neuron polysemanticity, which could be applied to better evaluate the efficacy of Sparse Auto-Encoders (SAEs).

可解释性神经元分析多义性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。