arXiv:2512.10978q-bio.NCcs.AI2025-12NeurIPS被引 5

通过分析注意力头的功能,揭示大模型推理的内在机制。

Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning

  • 构建新数据集CogQA,分解复杂问题为带思维链的子问题。
  • 发现注意力头具功能专一性,且在不同任务中分布稀疏而多样。
  • 移除或增强特定注意力头可显著影响推理表现,适用于模型优化。

大型语言模型(LLMs)在多种任务上达到顶尖性能,但其内部机制仍不透明。理解这些机制对提升其推理能力至关重要。受神经过程与人类认知相互作用的启发,我们提出一种新的可解释性框架,系统分析注意力头在LLMs中的角色与行为。我们引入了CogQA数据集,将复杂问题分解为具有思维链设计的逐步子问题,并关联到特定认知功能(如信息检索或逻辑推理)。通过多类探测方法,我们识别出负责这些功能的注意力头。在多个LLM家族中的分析表明,注意力头表现出功能特化,称为认知头。这些认知头具有普遍稀疏性,其数量和分布随不同认知功能而异,并呈现交互与层级结构。进一步实验显示,认知头在推理任务中起关键作用——移除它们导致性能下降,增强则提升推理准确率。这些发现深化了对LLM推理的理解,并对模型设计、训练与微调策略具有重要启示。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved state-of-the-art performance in a variety of tasks, but remain largely opaque in terms of their internal mechanisms. Understanding these mechanisms is crucial to improve their reasoning abilities. Drawing inspiration from the interplay between neural processes and human cognition, we propose a novel interpretability framework to systematically analyze the roles and behaviors of attention heads, which are key components of LLMs. We introduce CogQA, a dataset that decomposes complex questions into step-by-step subquestions with a chain-of-thought design, each associated with specific cognitive functions such as retrieval or logical reasoning. By applying a multi-class probing method, we identify the attention heads responsible for these functions. Our analysis across multiple LLM families reveals that attention heads exhibit functional specialization, characterized as cognitive heads. These cognitive heads exhibit several key properties: they are universally sparse, vary in number and distribution across different cognitive functions, and display interactive and hierarchical structures. We further show that cognitive heads play a vital role in reasoning tasks - removing them leads to performance degradation, while augmenting them enhances reasoning accuracy. These insights offer a deeper understanding of LLM reasoning and suggest important implications for model design, training, and fine-tuning strategies.

大模型注意力头可解释性推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。