arXiv:2604.08764cs.CLmath.DG2026-04被引 3

揭示语言模型训练中梯度方向的几何偏移机制

Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics

论文配图:Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
图 1 · 摘自论文原文
  • 通过激活数据构建低秩切线代理,追踪训练过程中的梯度方向
  • 发现激活导出方向携带更大梯度能量与显著更强的各向异性
  • 为梯度各向异性提供切线对齐的新解释,适合模型可解释性研究者

自引入以来,Transformer架构在自然语言处理领域占据主导地位。然而,近期研究揭示了这些模型中固有的各向异性现象,对其几何解释构成重大挑战。以往的理论研究很少基于底层表征几何。本文通过推导频率偏差采样如何削弱曲率可见性,并解释为何训练会优先放大切线方向,拓展了相关理论。我们采用基于概念的机制可解释性方法,在训练过程中而非仅事后分析,拟合由激活数据衍生的低秩切线代理,并将其与常规反向传播的真实梯度进行对比。在编码器型和解码器型语言模型中均发现,这些激活导出的方向比同秩的普通对照组捕获了更大的梯度能量和显著更高的梯度各向异性,为各向异性源于切线对齐提供了强有力的实证支持。

原文摘要 · Abstract (English)

Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inherent anisotropy phenomenon in these models, presenting a significant challenge to their geometric interpretation. Previous theoretical studies on this phenomenon are rarely grounded in the underlying representation geometry. In this paper, we extend them by deriving geometric arguments for how frequency-biased sampling attenuates curvature visibility and why training preferentially amplify tangent directions. Empirically, we then use concept-based mechanistic interpretability during training, rather than only post hoc, to fit activation-derived low-rank tangent proxies and test them against ordinary backpropagated true gradients. Across encoder-style and decoder-style language models, we find that these activation-derived directions capture both unusually large gradient energy and a substantially larger share of gradient anisotropy than matched-rank normal controls, providing strong empirical support for a tangent-aligned account of anisotropy.

Transformer梯度各向异性可解释性表示几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。