发现大模型中语义特征随训练、层数和规模的涌现规律
The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models
- 用稀疏自编码器分析模型激活,定位语义特征出现的位置与时间
- 不同领域存在明确的时间与规模阈值,特征在特定阶段突然出现
- 早期层特征会重新出现在深层,挑战传统表征演化认知
本文研究大型语言模型(LLM)中可解释的类别化特征的涌现现象,分析其在训练检查点(时间)、Transformer 层(空间)以及模型规模(尺度)上的行为。通过稀疏自编码器进行机制可解释性分析,我们识别出特定语义概念在神经激活中的出现时机与位置。结果表明,在多个领域中,特征涌现存在清晰的时间与规模阈值。值得注意的是,空间分析揭示了意外的语义再激活现象:早期层的特征会在后期层重新出现,挑战了关于 Transformer 模型表征动态的标准假设。
原文摘要 · Abstract (English)
This paper studies the emergence of interpretable categorical features within large language models (LLMs), analyzing their behavior across training checkpoints (time), transformer layers (space), and varying model sizes (scale). Using sparse autoencoders for mechanistic interpretability, we identify when and where specific semantic concepts emerge within neural activations. Results indicate clear temporal and scale-specific thresholds for feature emergence across multiple domains. Notably, spatial analysis reveals unexpected semantic reactivation, with early-layer features re-emerging at later layers, challenging standard assumptions about representational dynamics in transformer models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。