arXiv:2502.06809cs.LGcs.AI2025-02被引 8

提出按激活范围定位神经元,提升大模型可解释性与可控性。

Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution

  • 用激活范围替代离散神经元,实现更精准的概念定位
  • 实验证明范围干预比传统掩码减少80%以上的副作用
  • 适合需要精细控制大模型行为的研究者与工程师

大型语言模型中普遍存在的多义性现象破坏了离散神经元与概念的对应关系,严重影响模型可解释性与可控性。我们系统分析了基于编码器和解码器的多种大模型在不同数据集上的表现,发现即使对特定语义概念高度敏感的神经元也持续表现出多义性。重要的是,我们观察到:概念相关的神经元激活幅度呈现明显区分、通常呈高斯分布且重叠极小的规律。基于此,我们提出假设:通过解释和干预神经元的激活范围,可实现更精确的可解释性与针对性操控。为此,我们提出NeuronLens——一种基于激活范围的新型解释与操控框架,将概念归属定位到单个神经元内的激活区间。大量实证评估表明,基于范围的干预能有效操控目标概念,同时相比传统的神经元级掩码,对辅助概念及整体模型性能的负面影响显著降低。

原文摘要 · Abstract (English)

Pervasive polysemanticity in large language models (LLMs) undermines discrete neuron-concept attribution, posing a significant challenge for model interpretation and control. We systematically analyze both encoder and decoder based LLMs across diverse datasets, and observe that even highly salient neurons for specific semantic concepts consistently exhibit polysemantic behavior. Importantly, we uncover a consistent pattern: concept-conditioned activation magnitudes of neurons form distinct, often Gaussian-like distributions with minimal overlap. Building on this observation, we hypothesize that interpreting and intervening on concept-specific activation ranges can enable more precise interpretability and targeted manipulation in LLMs. To this end, we introduce NeuronLens, a novel range-based interpretation and manipulation framework, that localizes concept attribution to activation ranges within a neuron. Extensive empirical evaluations show that range-based interventions enable effective manipulation of target concepts while causing substantially less collateral degradation to auxiliary concepts and overall model performance compared to neuron-level masking.

可解释性大模型神经元干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。