用改进的局部学习系数分析注意力头如何随训练分化与专业化。
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
- 引入改进的局部学习系数(rLLC)量化注意力头的学习复杂度。
- 发现注意力头在训练中分化出不同功能角色,形成新型多字节电路。
- 为理解模型训练过程中的结构演化提供可量化的解释工具。
我们提出了一种基于奇异学习理论的改进型局部学习系数(rLLC),用于研究变换器语言模型在训练过程中内部结构的发展。通过将rLLC应用于两层仅注意力的变换器的各个组件,我们获得了关于注意力头逐步分化与专业化的全新见解。该方法揭示了注意力头如何在训练过程中演变为具有不同功能的角色,分析了这些头所专门处理的数据类型,并发现了一种此前未被识别的多字节电路。这些成果表明,rLLCs为发展性可解释性提供了原则性、量化的工具,旨在通过学习过程的演化来理解模型。更广泛地说,这项工作推进了数据分布结构、损失函数几何特性、学习动态与神经网络中涌现计算结构之间对应关系的研究。
原文摘要 · Abstract (English)
We introduce refined variants of the Local Learning Coefficient (LLC), a measure of model complexity grounded in singular learning theory, to study the development of internal structure in transformer language models during training. By applying these \textit{refined LLCs} (rLLCs) to individual components of a two-layer attention-only transformer, we gain novel insights into the progressive differentiation and specialization of attention heads. Our methodology reveals how attention heads differentiate into distinct functional roles over the course of training, analyzes the types of data these heads specialize to process, and discovers a previously unidentified multigram circuit. These findings demonstrate that rLLCs provide a principled, quantitative toolkit for \textit{developmental interpretability}, which aims to understand models through their evolution across the learning process. More broadly, this work takes a step towards establishing the correspondence between data distributional structure, geometric properties of the loss landscape, learning dynamics, and emergent computational structures in neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。