arXiv:2604.06256cs.LGcs.AI2026-04被引 2

通过频谱边缘分析揭示学习中的低维功能模式。

Spectral Edge Dynamics Reveal Functional Modes of Learning

  • 从参数更新方向识别主导频谱边缘,捕捉学习动态本质
  • 不同任务下频谱边缘呈现傅里叶模式、对数基模式等结构,提升学习集中度
  • 适合研究模型内在功能机制的学者,尤其关注任务对称性与学习路径者

训练过程中的突现学习(grokking)现象集中在少数主导更新方向——频谱边缘上,这些方向能可靠区分突现学习与非突现学习状态。我们发现标准可解释性工具(如头归因、激活探测、稀疏自编码器)无法捕捉这些方向:其结构在参数或特征空间中并不局域化。相反,每个方向在输入域上诱导出有结构的功能,揭示了表征层面分析无法发现的低维功能模式。对于模加法任务,所有主方向坍缩为单一傅里叶模式;对于乘法任务,仅在离散对数基下出现相同坍缩,带来5.9倍的集中度提升;对于减法任务,频谱边缘涵盖一个小的多模式族;对于 $x^2+y^2$,单一谐波基不足够,但加法与乘法特征的交叉项提供4倍方差增益,与 $(a+b)^2 - 2ab$ 分解一致。多任务训练强化了这种组合结构,$x^2+y^2$ 的频谱边缘继承加法电路的特征频率,集中度提升2.3倍。结果表明,训练发现了依赖于任务代数对称性的低维功能子空间。

原文摘要 · Abstract (English)

Training dynamics during grokking concentrate along a small number of dominant update directions -- the spectral edge -- which reliably distinguishes grokking from non-grokking regimes. We show that standard mechanistic interpretability tools (head attribution, activation probing, sparse autoencoders) fail to capture these directions: their structure is not localized in parameter or feature space. Instead, each direction induces a structured function over the input domain, revealing low-dimensional functional modes invisible to representation-level analysis. For modular addition, all leading directions collapse to a single Fourier mode. For multiplication, the same collapse appears only in the discrete-log basis, yielding a 5.9x improvement in concentration. For subtraction, the edge spans a small multi-mode family. For $x^2+y^2$, no single harmonic basis suffices, but cross-terms of additive and multiplicative features provide a 4x variance boost, consistent with the decomposition (a+b)^2 - 2ab. Multitask training amplifies this compositional structure, with the $x^2+y^2$ spectral edge inheriting the addition circuit's characteristic frequency (2.3x concentration increase). These results suggest that training discovers low-dimensional functional modes over the input domain, whose structure depends on the algebraic symmetry of the task. These results suggest that spectral edge dynamics identify low-dimensional functional subspaces governing learning, whose representation depends on the algebraic structure of the task. Simple harmonic structure emerges only when the task admits a symmetry-adapted basis; more complex tasks require richer functional descriptions.

机器学习可解释性学习动态频谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。