arXiv:2607.00603cs.LG2026-07

无需对齐即可精准识别网络中的死方向及其结构类型。

Measuring Dead Directions: Decomposing and Classifying Singular Structure off Canonical Alignment

  • 基于方向Fisher率直接读取每个死方向的阶数,不依赖优化路径。
  • 可区分真实奇异点与平庸对称性,准确恢复架构预测的阶数。
  • 适用于Transformer、卷积等多类层,适合研究模型内部结构的学者。

我们提出一种无需下降和对齐的奇异结构测量方法,在单个冻结检查点上,通过方向Fisher率精确恢复每个死方向的阶数 $k$,进而得到每方向的学习系数 $1/(2k)$,且该结果在优化器任意基下均成立。同一读取可分类方向:区分由架构固定的真奇异点与平庸的平坦规范对称性;当阶数无法确定时,方向Fisher幅值可补足判断。一个可插拔的检测器适用于Transformer、卷积及归一化层。该方法在构造单元和训练网络中均成功恢复架构预测的阶数,包括微调视觉Transformer中由LayerNorm核诱导的死结构,以及从头训练的压缩MLP在激活阶处形成的节点死亡现象。当奇异结构可枚举时,各方向阶数通过局部轨迹的类型交集,组装成全局系数 $(λ, m)$,符合闭式表达。本方法消除了底层速率结果所需的规范对齐和下降前提,将阶数恢复变为确定性的、架构通用的读取操作。进一步地,阶数决定普适奇异波动 $ν(k)$,但实际网络实现的 $ν$ 低于理论值,因活跃结构吸收了死方向的数据波动;多重性在单一局部假设下可从主导结构恢复。

原文摘要 · Abstract (English)

We give a descent-free, alignment-free measurement of singular structure on trained networks. At a single frozen checkpoint the read recovers the order $k$ of each dead direction from the directional-Fisher rate, the master invariant from which the per-direction learning coefficient $1/(2k)$ follows exactly, in whatever basis the optimizer left. The same read classifies each direction, separating a genuine singularity, whose order the architecture fixes, from a flat gauge symmetry; the directional-Fisher magnitude settles the cases the order cannot. A pluggable detector supplies the directions for transformer, convolutional, and normalisation layers. The read recovers the architecture-predicted order across constructed cells and trained networks, including a fine-tuned vision transformer whose dead structure is the LayerNorm-kernel gauge and a from-scratch one whose compressed MLP forms a node-death at its activation order. Where the singular structure enumerates, the per-direction orders assemble, through the typed intersection of the loci, into the global coefficient $(λ, m)$ matching the closed form. The method removes the canonical-alignment and descent preconditions of the underlying rate result, turning order-recovery into a deterministic, architecture-general reading. We then map its reach into the Watanabe triple: the order determines the universal singular fluctuation $ν(k)$, though a trained network's realized $ν$ falls below it as the live structure absorbs the dead direction's data fluctuation, and the multiplicity recovers from the dominant structure under a single-locus assumption.

神经网络奇异结构死方向模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。