arXiv:2601.04548cs.CLcs.AI2026-01被引 1

通过好坏神经元对比,揭示大模型任务决策机制。

Identifying Good and Bad Neurons for Task-Level Controllable LLMs

  • 基于生物拮抗原理,区分促进与抑制任务的神经元。
  • 在四大NLP任务中优于现有方法,准确率提升显著。
  • 适合研究模型可解释性与可控生成的学者使用。

大规模语言模型在多项选择题问答基准上表现出色,但其大规模神经元背后的复杂机制仍不清晰,给理解与调控模型带来挑战。现有研究虽能识别特定能力对应的神经元,但仅关注正向支持型神经元,忽略抑制性神经元,且易受模型偶然正确回答的误导(即巧合答对而非真正理解)。为此,我们提出NeuronLLM,一种基于功能拮抗原理的任务级语言模型理解框架。核心思想是:任务表现由两类对立神经元共同决定——促进任务完成的‘好神经元’和抑制任务完成的‘坏神经元’。通过对比学习建模好坏神经元,并利用增强问题集缓解模型的偶然行为,实现对神经元的全面刻画。在多种规模与架构的大模型上进行的综合实验表明,该方法在四个NLP任务中均显著优于现有方法,为理解大模型的功能组织提供了新视角。

原文摘要 · Abstract (English)

Large Language Models have demonstrated remarkable capabilities on multiple-choice question answering benchmarks, but the complex mechanisms underlying their large-scale neurons remain opaque, posing significant challenges for understanding and steering LLMs. While recent studies made progress on identifying responsible neurons for certain abilities, these ability-specific methods are infeasible for task-focused scenarios requiring coordinated use of multiple abilities. Moreover, these approaches focus only on supportive neurons that correlate positively with task completion, while neglecting neurons with other roles-such as inhibitive roles-and misled neuron attribution due to fortuitous behaviors in LLMs (i.e., correctly answer the questions by chance rather than genuine understanding). To address these challenges, we propose NeuronLLM, a novel task-level LLM understanding framework that adopts the biological principle of functional antagonism for LLM neuron identification. The key insight is that task performance is jointly determined by neurons with two opposing roles: good neurons that facilitate task completion and bad neurons that inhibit it. NeuronLLM achieves a holistic modeling of neurons via contrastive learning of good and bad neurons, while leveraging augmented question sets to mitigate the fortuitous behaviors in LLMs. Comprehensive experiments on LLMs of different sizes and families show the superiority of NeuronLLM over existing methods in four NLP tasks, providing new insights into LLM functional organization.

可解释性神经元分析大模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。