找出大模型中完成语法任务的关键注意力头,揭示其结构与分工。
$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks
- 提出新方法K-MSHC,定位完成特定任务的最小关键注意力头集合。
- 发现语法任务依赖浅层头,算术问题激活深浅层,验证任务分布更分散。
- 不同任务共享弱头或强头,但各有专用'超级头',体现能力专精与复用并存。
理解中等规模语言模型(≤100亿参数)中驱动特定能力的神经组件仍是一大挑战。我们提出$(\bm{K}, ε)$-最小充分头电路(K-MSHC)方法,用于识别分类任务中的关键注意力头,并设计高效算法Search-K-MSHC以发现这些电路。将Search-K-MSHC应用于Gemma-9B,分析三类句法任务:语法可接受性、算术验证和算术应用题。结果表明,各类任务有特定头电路:语法任务主要使用早期层头,算术应用题在浅层与深层均有显著活动,算术验证则呈现更分布式模式。发现非线性重叠规律:语法与算术共享多个“弱”头,而算术与应用题更共用关键“强”头。重要的是,每类任务均有独立的“超级头”,跨任务重叠极小,说明句法与数值能力源于专用但部分可复用的头电路。
原文摘要 · Abstract (English)
Understanding which neural components drive specific capabilities in mid-sized language models ($\leq$10B parameters) remains a key challenge. We introduce the $(\bm{K}, ε)$-Minimum Sufficient Head Circuit ($K$-MSHC), a methodology to identify minimal sets of attention heads crucial for classification tasks as well as Search-K-MSHC, an efficient algorithm for discovering these circuits. Applying our Search-K-MSHC algorithm to Gemma-9B, we analyze three syntactic task families: grammar acceptability, arithmetic verification, and arithmetic word problems. Our findings reveal distinct task-specific head circuits, with grammar tasks predominantly utilizing early layers, word problems showing pronounced activity in both shallow and deep regions, and arithmetic verification demonstrating a more distributed pattern across the network. We discover non-linear circuit overlap patterns, where different task pairs share computational components at varying levels of importance. While grammar and arithmetic share many "weak" heads, arithmetic and word problems share more consistently critical "strong" heads. Importantly, we find that each task maintains dedicated "super-heads" with minimal cross-task overlap, suggesting that syntactic and numerical competencies emerge from specialized yet partially reusable head circuits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。