arXiv:2511.21610cs.CL2025-11

通过辅助指标定位大模型中编码特定技能的神经元,无需人工标注即可发现隐藏推理捷径。

Auxiliary Metrics Help Decoding Skill Neurons in the Wild

  • 用外部标签和置信度等辅助指标关联神经元激活
  • 在生成与推理任务中识别出已知技能及新发现的算术捷径
  • 方法轻量通用,适合研究模型内部可解释性

大型语言模型在众多任务中表现出色,但其内部机制仍不透明。本文提出一种简单、轻量且普适的方法,旨在分离编码特定技能的神经元。基于先前通过分类任务软提示训练识别‘技能神经元’的工作,本方法将分析扩展至包含多种技能的复杂场景。通过将神经元激活与外部标签、模型自身置信度等辅助指标相关联,无需手动词元聚合即可揭示可解释且任务特定的行为。我们在开放式文本生成和自然语言推理任务上验证了该方法,结果表明其不仅能检测出已知技能相关的神经元,还能发现BigBench算术推理中的此前未被识别的捷径。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit remarkable capabilities across a wide range of tasks, yet their internal mechanisms remain largely opaque. In this paper, we introduce a simple, lightweight, and broadly applicable method with a focus on isolating neurons that encode specific skills. Building upon prior work that identified "skill neurons" via soft prompt training on classification tasks, our approach extends the analysis to complex scenarios involving multiple skills. We correlate neuron activations with auxiliary metrics -- such as external labels and the model's own confidence score -- thereby uncovering interpretable and task-specific behaviors without the need for manual token aggregation. We empirically validate our method on tasks spanning open-ended text generation and natural language inference, demonstrating its ability to detect neurons that not only drive known skills but also reveal previously unidentified shortcuts in arithmetic reasoning on BigBench.

可解释性技能神经元大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。