拆解神经网络权重的大小与方向,发现方向决定解决方案,大小影响稳定性。
Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking
- 将不同训练路径的权重分解为大小和方向,交叉重组后继续训练。
- 方向可迁移并决定最终解,40/40次实验中成功引导至捐赠者解。
- 适合研究模型内部机制或可解释性的研究人员阅读。
部分训练的神经网络中哪些属性具有因果可移植性?单轨迹干预仅显示某次运行内的必要性,而非跨运行可移植性。我们提出跨轨迹嵌合体干预:给定两个不同随机种子的训练结果,将每个权重向量拆分为模长和单位方向,将一个运行的模长与另一个的单位方向重组后继续训练。在两个发生‘领悟’(grokking)的模运算任务上,这两者可分离。方向携带可移植、捐赠者特异的电路身份:将捐赠者的方向植入接收者的模长,40/40次实验均引导该运行走向捐赠者的电路;而角度匹配的随机控制组无此效果。该转移呈阈值特性,其位置由接收者的模长预测,在全部20对组合中按模长类别完全分离(联合置换概率1.9e-4)。模长仅带来微弱、分布式的延迟效应,无身份信号。自适应二分法将阈值定位至±1/64范围内。方向决定轨迹趋近哪个解,模长则控制该身份被覆盖的敏感程度。
原文摘要 · Abstract (English)
Which properties of a partially trained network are causally portable to a different, independently trained network? Single-trajectory interventions show necessity within one run, not portability across runs. We introduce cross-trajectory chimera interventions: given two runs from different seeds, we split each weight vector into a norm and a unit direction, recombine one run's norm with the other's direction, and continue training. On two modular-arithmetic tasks that grok, the components dissociate. Direction carries a transferable, donor-specific circuit identity: implanting a donor's direction at the recipient's norm drives the run to the donor's circuit in 40/40 cases, while an angle-matched random control yields no shift. The transfer is threshold-like, and its location is predicted by the recipient's norm, separating perfectly by norm class over all 20 pairs (joint permutation probability 1.9e-4). Norm carries only a modest, distributed delay effect and no identity signal. An adaptive bisection procedure localizes the threshold to +/-1/64. Direction indexes which solution a trajectory approaches; norm governs how susceptible that identity is to being overwritten.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。