解析大模型注意力头机制,揭示其推理逻辑
Attention Heads of Large Language Models: A Survey
- 构建四阶段思维框架,系统分析注意力头功能
- 分类梳理无模型与有模型检测方法,总结评估基准
- 适合关注模型可解释性与内部机理的研究者
自ChatGPT问世以来,大语言模型(LLMs)在各类任务中表现卓越,但其内部运作仍如黑箱一般。理解LLMs的推理瓶颈已成为关键挑战,因这些限制与其架构深度相关。其中,注意力头成为研究模型内在机制的核心焦点。本文通过构建受人类思维过程启发的四阶段框架——知识回忆、上下文识别、潜在推理和表达准备,系统梳理注意力头的作用与机制。基于此框架,全面回顾现有研究,识别并分类特定注意力头的功能。同时,分析发现这些特殊头的实验方法,将其分为无模型与需建模两类,并总结相关评估方法与基准。最后,讨论当前研究局限,并提出未来可能方向。
原文摘要 · Abstract (English)
Since the advent of ChatGPT, Large Language Models (LLMs) have excelled in various tasks but remain as black-box systems. Understanding the reasoning bottlenecks of LLMs has become a critical challenge, as these limitations are deeply tied to their internal architecture. Among these, attention heads have emerged as a focal point for investigating the underlying mechanics of LLMs. In this survey, we aim to demystify the internal reasoning processes of LLMs by systematically exploring the roles and mechanisms of attention heads. We first introduce a novel four-stage framework inspired by the human thought process: Knowledge Recalling, In-Context Identification, Latent Reasoning, and Expression Preparation. Using this framework, we comprehensively review existing research to identify and categorize the functions of specific attention heads. Additionally, we analyze the experimental methodologies used to discover these special heads, dividing them into two categories: Modeling-Free and Modeling-Required methods. We further summarize relevant evaluation methods and benchmarks. Finally, we discuss the limitations of current research and propose several potential future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。