统一分析模型、数据与训练过程,揭示行为背后的多维度机制
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
- 将AdamW训练的模型重构为核机器,实现组件/数据/训练轨迹的联合建模
- 提出参数级与步骤级影响分数,在多个场景中验证了行为分解的有效性
- 适用于理解复杂模型的学习阶段,如Transformer的泛化与语言模型预训练
事后可解释性方法通常孤立地将模型行为归因于组件、数据或训练轨迹,且局限于局部到全局的某一粒度。为此,我们提出ExPLAIND,一个理论严谨、统一的框架,整合模型组件、数据与训练轨迹,并支持跨粒度解释。我们将使用AdamW训练的模型重新表述为核机器,从所得核特征映射中推导出新型参数级和步骤级影响分数。在多个设置中实证验证了模型行为的分解效果,并应用于两个案例研究:对展现Grokking现象的Transformer,支持既有学习阶段划分,并将最终阶段修正为外层围绕记忆后习得的表示管道对齐;在EuroLLM预训练中,揭示两阶段动态——第一阶段外层MLP学习为主,第二阶段中间注意力层相对影响增强。结果确立ExPLAIND作为统一解释模型行为与训练动态的框架。
原文摘要 · Abstract (English)
Post-hoc interpretability methods typically attribute a model's behavior to its components, data, or training trajectory in isolation, and are often tied to a particular level of granularity along the local-to-global spectrum. This leads to explanations that lack a unified view and may miss key interactions. We present ExPLAIND, a theoretically grounded, unified framework that integrates model components, data, and training trajectory while supporting explanations across granularities. We generalize recent work on gradient path kernels, reformulating models trained by AdamW as kernel machines. From the resulting kernel feature maps, we derive novel parameter-wise and step-wise influence scores. We empirically validate the resulting decomposition of model behavior in several settings and apply ExPLAIND to two case studies. Our findings on a Transformer exhibiting Grokking support previously proposed learning phases, while refining the final phase as one in which outer layers align around a representation pipeline learned after memorization. For EuroLLM pretraining, ExPLAIND reveals a two-phase dynamic, with the first characterized by outer-layer MLP learning and the second by increased relative influence of intermediate attention layers. These results establish ExPLAIND as a unified framework for interpreting model behavior and training dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。