arXiv:2604.00443cs.CLcs.AI2026-04

发现神经元激活重叠常由词汇同形导致,而非概念压缩。

Polysemanticity or Polysemy? Lexical Identity Confounds Superposition Metrics

  • 通过2×2分解实验分离词汇与语义因素,量化混淆来源。
  • 词汇同形导致的激活重叠在多数模型中高于语义相似性。
  • 去除该混淆可提升词义消歧与知识编辑效果。

若同一神经元对'lender'和'riverside'均有响应,传统指标会归因于超叠加——即压缩了两个无关概念。本文探究这种重叠有多少源于词汇混淆:神经元对共享词形(如'bank')响应,而非两个压缩概念。2×2因子分解显示,在110M至70B参数量的多个模型中,仅由词汇相同但意义不同引起的激活重叠,始终超过由词形不同但意义相同引起的重叠。该混淆现象也存在于稀疏自编码器中(18-36%的特征混合不同词义),且仅占据≤1%的激活维度,还损害下游任务表现:过滤后词义消歧性能提升,知识编辑更精准(p=0.002)。

原文摘要 · Abstract (English)

If the same neuron activates for both "lender" and "riverside," standard metrics attribute the overlap to superposition--the neuron must be compressing two unrelated concepts. This work explores how much of the overlap is due a lexical confound: neurons fire for a shared word form (such as "bank") rather than for two compressed concepts. A 2x2 factorial decomposition reveals that the lexical-only condition (same word, different meaning) consistently exceeds the semantic-only condition (different word, same meaning) across models spanning 110M-70B parameters. The confound carries into sparse autoencoders (18-36% of features blend senses), sits in <=1% of activation dimensions, and hurts downstream tasks: filtering it out improves word sense disambiguation and makes knowledge edits more selective (p = 0.002).

神经元解释词汇混淆模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。