arXiv:2509.24164cs.CL2025-09被引 1

通过注意力头分析,揭示大模型如何在上下文学习中识别任务并学习解法。

Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis

  • 基于任务子空间逻辑归因,定位负责任务识别与学习的特定注意力头。
  • 发现任务识别头对齐隐藏状态,学习头则在子空间内旋转至正确标签。
  • 可解释多种已有发现,适合关注模型机制可解释性的研究者。

我们通过融合注意力头层面的组件分析与整体分解方法,研究大语言模型中上下文学习(ICL)的机制。提出基于任务子空间逻辑归因(TSLA)的新框架,识别出专门负责任务识别(TR)和任务学习(TL)的注意力头,并证明它们独立而有效。通过相关性分析、消融实验与输入扰动验证,这些头部分别捕获ICL中的相应成分。借助隐藏状态的几何操控实验,发现任务识别头通过将隐藏状态对齐到任务子空间来促进识别,而任务学习头则在子空间内旋转状态以指向正确标签。该框架统一解释了以往关于诱导头与任务向量等发现,为大模型在多任务场景下的上下文学习提供了可解释的统一机制。

原文摘要 · Abstract (English)

We investigate the mechanistic underpinnings of in-context learning (ICL) in large language models by reconciling two dominant perspectives: the component-level analysis of attention heads and the holistic decomposition of ICL into Task Recognition (TR) and Task Learning (TL). We propose a novel framework based on Task Subspace Logit Attribution (TSLA) to identify attention heads specialized in TR and TL, and demonstrate their distinct yet complementary roles. Through correlation analysis, ablation studies, and input perturbations, we show that the identified TR and TL heads independently and effectively capture the TR and TL components of ICL. Using steering experiments with geometric analysis of hidden states, we reveal that TR heads promote task recognition by aligning hidden states with the task subspace, while TL heads rotate hidden states within the subspace toward the correct label to facilitate prediction. We further show how previous findings on ICL mechanisms, including induction heads and task vectors, can be reconciled with our attention-head-level analysis of the TR-TL decomposition. Our framework thus provides a unified and interpretable account of how large language models execute ICL across diverse tasks and settings.

上下文学习注意力头分析可解释性大模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。