arXiv:2509.24198cs.LG2025-09被引 2

发现负激活在语法理解中起关键作用,突破传统只关注正激活的偏见。

Negative Pre-activations Differentiate Syntax

  • 聚焦罕见的Wasserstein神经元,发现其负预激活具主动功能而非副作用。
  • 仅屏蔽少数此类神经元的负值,语法性能骤降,困惑度上升20%以上。
  • 适用于研究语言模型内部机制与语法表征的开发者和研究员。

现代大语言模型广泛采用GELU或SiLU等平滑激活函数,使负预激活可携带信号与梯度。然而,许多神经元级可解释性分析仍聚焦于大正激活,隐含认为负区不重要——这延续了ReLU时代的认知。本文挑战此假设,探究负预激活是否被模型利用。通过研究输出分布显著偏离高斯基线的稀疏子集Wasserstein神经元,我们发现该负区域并非单纯梯度优化副产品,而是具有功能性。仅对少量此类神经元的负预激活进行零值干预,便导致困惑度显著上升,并在BLiMP和TSE上严重损害语法性能;而随机或困惑度匹配的多神经元负激活移除则基本不影响语法表现。相反,在非语法基准上,控制组(困惑度匹配)的破坏性更强,形成语法与非语法能力的双重分离。词性分析表明,异常信息集中于句法骨架词元,层间干预显示局部退化逐层累积,训练动态分析揭示该干预危害随Wasserstein神经元出现与稳定而增强。结果表明,光滑激活模型中,少数Wasserstein神经元的负预激活是语法处理的关键主动成分。

原文摘要 · Abstract (English)

Modern large language models increasingly use smooth activation functions such as GELU or SiLU, allowing negative pre-activations to carry both signal and gradient. Nevertheless, many neuron-level interpretability analyses have historically focused on large positive activations, often implicitly treating the negative region as less informative, a carryover from the ReLU-era. We challenge this assumption and ask whether and how negative pre-activations are leveraged by models. We address this question by studying a sparse subpopulation of Wasserstein neurons whose output distributions deviate strongly from a Gaussian baseline and that functionally differentiate similar inputs. We show that this negative region plays an active role rather than reflecting a mere gradient optimization side effect. A minimal, sign-specific intervention that zeroes only the negative pre-activations of a small set of Wasserstein neurons substantially increases perplexity and sharply degrades grammatical performance on BLiMP and TSE, whereas both random and perplexity-matched ablations of many more non-Wasserstein neurons in their negative pre-activations leave grammatical performance largely intact. Conversely, on a suite of non-grammatical benchmarks, the perplexity-matched control ablation is more damaging than the Wasserstein neuron ablation, yielding a double dissociation between syntax and other capabilities. Part-of-speech analysis localizes the excess surprisal to syntactic scaffolding tokens, layer-specific interventions show that small local degradations accumulate across depth, and training-dynamics analysis reveals that the same sign-specific ablation becomes more harmful as Wasserstein neurons emerge and stabilize. Together, these results identify negative pre-activations in a sparse subpopulation of Wasserstein neurons as an actively used substrate for syntax in smooth-activation language models.

语言模型语法理解神经元分析负激活

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。