揭示语言模型中多种机制现象的统一根源。
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
- 从数据生成的分层潜在结构出发,解释模型现象
- 在合成与真实语料上验证理论有效性
- 适合关注模型可解释性的研究者阅读
当前关于机制可解释性的研究发现了基于Transformer的语言模型中诸多令人困惑的现象,如归纳头、函数向量和水蛇效应。这些现象曾被分别关联到不同的数据分布特性或模型架构设计,但缺乏对数据、模型架构与优化三者关系的统一理解。本文提出,这些现象均源于数据生成过程中的分层潜在结构,结合加性模块间梯度去相关性及表征几何的方向凹性。我们在小规模模拟模型和大规模合成数据设置下验证了该理论,并与自然语言数据训练的语言模型进行对比,结果支持其普适性。
原文摘要 · Abstract (English)
Contemporary studies in mechanistic interpretability have uncovered many puzzling phenomena in the neural information processing of Transformer-based language models, such as induction heads, function vectors, and the Hydra effect. Some of these individual phenomena have been independently tied to different data distributional properties, while some have been loosely associated with model architecture and how Transformers process information. However, a unified understanding of the relationship between data, model architecture, and optimization remains lacking, failing to answer the fundamental question: why do these three phenomena appear universally across different model families and scales, despite their seeming disconnect? In this work, we answer this question by unifying these three phenomena as consequences of hierarchical latent structures in the data generation process, coupled with decorrelated gradients across additive model components and directional concavity in the representation geometry. We validate our theoretical results in a toy model regime and in a large-scale synthetic data regime, comparing them with language models trained on natural language data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。