arXiv:2505.21191cs.CLcs.LG2025-05被引 3

揭示大模型执行指令的神经机制,找到关键稀疏组件。

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities

  • 提出SPARCOM框架,定位模型中特定指令的稀疏神经元与专家。
  • 发现这些组件具有功能通用性和唯一性,对指令执行至关重要。
  • 适用于研究模型可解释性、安全对齐与高效微调的科研人员。

大型语言模型(LLM)的微调显著提升了其指令遵循能力,但其背后的计算机制仍不清晰。本研究系统分析了微调如何重构模型计算,通过隔离并分析指令特定的稀疏组件——即密集模型中的神经元,以及混合专家(MoE)架构中的神经元和专家。为此,我们构建了涵盖六个不同类别的平衡指令数据集HexaInst,并提出SPARCOM分析框架,包含三项核心贡献:(1)识别这些稀疏组件的方法;(2)评估其功能通用性与独特性;(3)系统比较其变化。实验表明,这些组件具备功能通用性、独特性,并在指令执行中起关键作用。本工作揭示了微调引发的适应性与稀疏计算基底之间的关系,为可信大模型社区提供了更深层的理解。

原文摘要 · Abstract (English)

The finetuning of Large Language Models (LLMs) has significantly advanced their instruction-following capabilities, yet the underlying computational mechanisms driving these improvements remain poorly understood. This study systematically examines how fine-tuning reconfigures LLM computations by isolating and analyzing instruction-specific sparse components, i.e., neurons in dense models and both neurons and experts in Mixture-of-Experts (MoE) architectures. In particular, we introduce HexaInst, a carefully curated and balanced instructional dataset spanning six distinct categories, and propose SPARCOM, a novel analytical framework comprising three key contributions: (1) a method for identifying these sparse components, (2) an evaluation of their functional generality and uniqueness, and (3) a systematic comparison of their alterations. Through experiments, we demonstrate functional generality, uniqueness, and the critical role of these components in instruction execution. By elucidating the relationship between fine-tuning-induced adaptations and sparse computational substrates, this work provides deeper insights into how LLMs internalize instruction-following behavior for the trustworthy LLM community.

大模型可解释性指令遵循稀疏激活MoE架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。