arXiv:2510.00468cs.LGcs.AI2025-10被引 3

通过分析eNTK特征值,发现训练后模型中的关键特征方向。

Feature Identification via the Empirical NTK

  • 用eNTK特征分析揭示模型内部特征方向
  • 在模运算任务中对齐傅里叶特征,语言模型中优于PCA基线
  • 适合关注模型可解释性的研究者

我们证明,对经验神经正切核(eNTK)进行特征分解可以揭示训练后神经网络中的特征方向。在三个逐步接近现实的设置中——一个在模加法任务上训练的一层MLP、一个在模加法任务上训练的一层Transformer,以及预训练语言模型Gemma-3-270M——我们发现eNTK的前几大特征空间与真实或可解释的特征对齐。在模算术例子中,eNTK的主特征空间与MLP使用的傅里叶特征,以及Transformer在特定种子频率下用于实现已知算法的傅里叶特征对齐。此外,相关子空间的对齐随训练演化,在‘领悟’现象开始时,其一阶导数达到峰值。对于Gemma-3-270M,我们在TinyStories上下文窗口数据集上计算了eNTK的前几个特征方向,并检验其与自动生成的词性及其他语法特征方向的对齐情况。结果显示,eNTK特征方向在语法特征上的对齐性能优于同等预算的模型激活上进行PCA的基线方法。这些结果表明,eNTK特征分析可能为机械可解释性提供一种新途径。

原文摘要 · Abstract (English)

We provide evidence that eigenanalysis of the empirical neural tangent kernel (eNTK) can surface feature directions in trained neural networks. Across three increasingly realistic settings -- a 1-layer MLP trained on modular addition, a 1-layer Transformer trained on modular addition and the pretrained language model Gemma-3-270M -- we show that top eigenspaces of the eNTK align with ground-truth or interpretable features. In the modular arithmetic examples, top eNTK eigenspaces align with the Fourier features used by the MLP and the Fourier features at seed-dependent frequencies used by the Transformer to implement known ground-truth algorithms. Moreover, the alignment of the relevant subspaces evolves over training, with its first derivative peaking near the onset of grokking. For Gemma-3-270M, we compute top eNTK eigendirections on a dataset of TinyStories context windows and check their alignment with an automatically-generated set of parts-of-speech and other grammatical feature directions. We find that the alignment of eNTK eigendirections with grammar features outperforms a same-budget baseline of PCA on model activations. These results suggest that eNTK eigenanalysis may provide a new handle towards identifying features in trained models for mechanistic interpretability.

可解释性eNTK特征识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。