通过追踪特征演变,揭示大模型预训练的内在规律
Evolution of Concepts in Language Model Pre-Training
- 用跨编码器方法解析模型各阶段的可解释特征
- 发现特征在中期开始形成,后期出现复杂模式
- 为理解模型学习动态提供细粒度观测视角
语言模型通过预训练获得强大能力,但该过程仍如黑箱。本文使用一种称为 crosscoders 的稀疏字典学习方法,追踪预训练快照中线性可解释特征的演化。我们发现,大多数特征在特定阶段开始形成,而更复杂的模式在后期出现。特征归因分析揭示了特征演化与下游性能之间的因果关系。我们的特征级观察与先前关于 Transformer 两阶段学习过程的研究高度一致,我们将其称为统计学习阶段和特征学习阶段。本工作为追踪语言模型学习动态中的细粒度表征进展打开了可能。
原文摘要 · Abstract (English)
Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-training snapshots using a sparse dictionary learning method called crosscoders. We find that most features begin to form around a specific point, while more complex patterns emerge in later training stages. Feature attribution analyses reveal causal connections between feature evolution and downstream performance. Our feature-level observations are highly consistent with previous findings on Transformer's two-stage learning process, which we term a statistical learning phase and a feature learning phase. Our work opens up the possibility to track fine-grained representation progress during language model learning dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。