arXiv:2412.18053cs.CLcs.AI2024-12ACL被引 3

发现神经元激活与输出间存在全局线性关系,可量化其控制能力。

Neuron Empirical Gradient: Discovering and Quantifying Neurons Global Linear Controllability

  • 通过干预神经元激活,发现其与输出呈全局线性关系。
  • 在MCEval8k上验证,该关系能有效表征模型知识。
  • 方法高效鲁棒,适合大规模神经元行为分析。

预训练语言模型中的前馈神经元虽可编码知识,但以往研究仅关注少数影响输出的神经元,导致整体神经元作用不清晰,限制了知识编辑等任务进展。本文在知识探测数据集上通过神经元干预,揭示神经元激活与输出间存在全局线性关系。该关系的梯度称为神经元经验梯度(NEG),用于量化激活变化对预测的影响。为高效计算NEG,提出NeurGrad,支持在大规模语言模型中进行神经元行为分析。通过技能神经元探测,证实NEG能有效捕捉多种提示下的语言能力。在多体裁多选题知识基准MCEval8k上的实验验证了NEG表示模型知识的能力。进一步分析表明,基于NEG的技能表征具有高效、鲁棒、灵活及相互依赖等特性。代码与数据已公开。

原文摘要 · Abstract (English)

While feed-forward neurons in pre-trained language models (PLMs) can encode knowledge, past research targeted a small subset of neurons that heavily influence outputs. This leaves the broader role of neuron activations unclear, limiting progress in areas like knowledge editing. We uncover a global linear relationship between neuron activations and outputs using neuron interventions on a knowledge probing dataset. The gradient of this linear relationship, which we call the neuron empirical gradient (NEG), captures how changes in activations affect predictions. To compute NEG efficiently, we propose NeurGrad, enabling large-scale analysis of neuron behavior in PLMs. We also show that NEG effectively captures language skills across diverse prompts through skill neuron probing. Experiments on MCEval8k, a multi-genre multiple-choice knowledge benchmark, support NEG's ability to represent model knowledge. Further analysis highlights the key properties of NEG-based skill representation: efficiency, robustness, flexibility, and interdependency. The code and data are released.

神经元分析知识表征语言模型可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。