通过追踪训练数据,揭示大模型可解释单元的形成原因。
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
- 用影响函数定位关键训练样本,反推可解释单元来源。
- 删减高影响力样本可显著改变模型可解释头的出现。
- 适合研究模型可解释性与训练数据关系的学者。
尽管机制可解释性已识别出大语言模型中的可解释电路,但其在训练数据中的因果起源仍不明确。我们提出机制数据溯源(Mechanistic Data Attribution, MDA),一种基于影响函数的可扩展框架,用于将可解释单元追溯至特定训练样本。在Pythia系列模型上的大量实验表明,针对性干预——移除或增强少量高影响力样本——会显著调控可解释头的出现,而随机干预则无效果。分析显示,重复性结构化数据(如LaTeX、XML)起到了机制催化作用。此外,针对归纳头形成的干预同时改变了模型的上下文学习能力,为归纳头与上下文学习之间长期存在的功能关联提供了直接因果证据。最后,我们提出一种机制数据增强流程,能一致加速不同规模模型的电路收敛,为引导大模型发展轨迹提供了一种原则性方法。
原文摘要 · Abstract (English)
While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive. We introduce Mechanistic Data Attribution (MDA), a scalable framework that employs Influence Functions to trace interpretable units back to specific training samples. Through extensive experiments on the Pythia family, we causally validate that targeted intervention--removing or augmenting a small fraction of high-influence samples--significantly modulates the emergence of interpretable heads, whereas random interventions show no effect. Our analysis reveals that repetitive structural data (e.g., LaTeX, XML) acts as a mechanistic catalyst. Furthermore, we observe that interventions targeting induction head formation induce a concurrent change in the model's in-context learning (ICL) capability. This provides direct causal evidence for the long-standing hypothesis regarding the functional link between induction heads and ICL. Finally, we propose a mechanistic data augmentation pipeline that consistently accelerates circuit convergence across model scales, providing a principled methodology for steering the developmental trajectories of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。