arXiv:2505.18752cs.CL2025-05NeurIPS被引 6

揭示大模型少样本学习中注意力头与任务向量的统一机制

Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning

  • 通过隐藏状态几何分析,发现早期层增强可分性,后期层提升对齐度
  • 实验证明前词注意力头负责可分性,归纳头和任务向量促进对齐
  • 为分类任务的上下文学习提供统一解释,适合研究模型内部机理者

上下文学习(ICL)的特殊性质引发了对大语言模型内部机制的研究。以往工作多聚焦于特定层的注意力头或任务向量,缺乏将这些组件与隐藏状态跨层演变关联的统一框架。本文针对分类任务提出此类框架,分析影响性能的两个几何因素:查询隐藏状态的可分性与对齐度。分层动态分析揭示显著的两阶段机制:早期层生成可分性,后期层实现对齐度提升。消融实验表明,前词注意力头驱动可分性,而归纳头和任务向量增强对齐度。研究弥合了注意力头与任务向量之间的鸿沟,为ICL提供了统一的内在机制解释。

原文摘要 · Abstract (English)

The unusual properties of in-context learning (ICL) have prompted investigations into the internal mechanisms of large language models. Prior work typically focuses on either special attention heads or task vectors at specific layers, but lacks a unified framework linking these components to the evolution of hidden states across layers that ultimately produce the model's output. In this paper, we propose such a framework for ICL in classification tasks by analyzing two geometric factors that govern performance: the separability and alignment of query hidden states. A fine-grained analysis of layer-wise dynamics reveals a striking two-stage mechanism: separability emerges in early layers, while alignment develops in later layers. Ablation studies further show that Previous Token Heads drive separability, while Induction Heads and task vectors enhance alignment. Our findings thus bridge the gap between attention heads and task vectors, offering a unified account of ICL's underlying mechanisms.

上下文学习注意力机制隐藏状态大模型机理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。