arXiv:2412.17626cs.LGcs.CL2024-12被引 14

通过持续追踪特征演化,揭示大模型训练中的机制规律。

Tracking the Feature Dynamics in LLM Training: A Mechanistic Study

  • 提出SAE-Track方法,连续获取稀疏自编码器以跟踪特征变化。
  • 发现特征在训练中呈现语义演进、形成过程与向量方向漂移。
  • 适合关注模型内部机制与训练动态的研究者参考。

理解训练动态与特征演化对大型语言模型(LLMs)的机械可解释性至关重要。尽管稀疏自编码器(SAEs)已被用于识别模型内部特征,但这些特征在训练过程中的演变仍不清晰。本研究提出(1)SAE-Track,一种高效获取连续系列SAEs的新方法,为机制研究奠定基础;(2)分析特征的语义演化;(3)揭示特征形成的内在过程;(4)观察特征向量的方向漂移。研究成果深化了对大模型特征动态的理解,有助于揭示训练机制与演化规律。代码已开源:https://github.com/Superposition09m/SAE-Track。

原文摘要 · Abstract (English)

Understanding training dynamics and feature evolution is crucial for the mechanistic interpretability of large language models (LLMs). Although sparse autoencoders (SAEs) have been used to identify features within LLMs, a clear picture of how these features evolve during training remains elusive. In this study, we (1) introduce SAE-Track, a novel method for efficiently obtaining a continual series of SAEs, providing the foundation for a mechanistic study that covers (2) the semantic evolution of features, (3) the underlying processes of feature formation, and (4) the directional drift of feature vectors. Our work provides new insights into the dynamics of features in LLMs, enhancing our understanding of training mechanisms and feature evolution. For reproducibility, our code is available at https://github.com/Superposition09m/SAE-Track.

特征演化可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。