arXiv:2606.21158cs.LGstat.ML2026-06被引 1

提出一种低成本方法,通过激活和梯度谱分析网络的奇异结构。

Dead-Direction Signatures: A Cheap Spectral Reading of Singular Complexity

  • 利用激活与梯度谱的谱特征,闭式计算网络的死方向。
  • 在秩缺陷为1~4时,斜率比值与理论预测2:3:4一致,精度达99%。
  • 适用于模型秩追踪,尤其在低预算下替代传统后验采样方法。

奇异学习理论通过损失函数奇点的几何结构刻画深度网络复杂性。标准的局部学习系数(LLC)通过SGLD估算Watanabe的实对数规范阈值(RLCT, $λ$),需每任务校准且每检查点耗时 $10^4$-$10^6$ 次前向-反向传播。本文提出死方向谱(DDS),一种基于闭式谱读取的廉价方法:从指定层的激活矩阵或样本梯度Fisher-Gram中读取谱特征,以谱线性代数替代SGLD后验链。该方法基于死方向框架,预测任意奇点处激活与Fisher谱间存在结构性相关,并引入秩乘体积恒等式,使单特征值监控无法实现的活跃体积 $\ ext{logdet}^{+}(G)$ 斜率可计数死方向,精确追踪 $r \in \{1,2,3,4\}$ 的秩缺陷,对应斜率比值分别为 $2.0, 3.1, 4.0$(理论值 $2,3,4$),最小特征值对秩不敏感。在降秩回归中,经校准的LLC恢复 $λ$ 均值达 $99\\$,而DDS可观测量按框架预测符号追踪 $λ$;在非线性模块加法变换器中,DDS在十八个数量级上分离 $d_{\mathrm{model}}$,而校准后的LLC在协议预算下呈现秩平坦。相较于LLC的集成后验读取,DDS提供方向性、层局部的死方向读取,闭式地从激活与梯度谱中获得。

原文摘要 · Abstract (English)

Singular learning theory characterises the complexity of a deep network through the geometry of its loss singularities. The local learning coefficient (LLC), the standard estimator of Watanabe's real log canonical threshold (RLCT, $λ$), reads this geometry as an integrated Bayesian scalar through SGLD, which needs per-task calibration and $10^4$-$10^6$ forward-backward passes per checkpoint. We introduce Dead-Direction Signatures (DDS), a family of cheap closed-form spectral readings of singular structure: each reads a network's activation matrix or per-sample-gradient Fisher-Gram at a chosen layer, replacing the SGLD posterior chain with spectral linear algebra. The readings rest on a dead-direction framework that predicts a structural correlation between activation- and Fisher-side spectra at any singular minimum, and a rank-multiplicative volume identity that single-eigenvalue monitors cannot produce: the active-volume $\log\det^{+}(G)$ slope counts the dead directions, tracking the rank-deficit $r$ across $r \in \{1,2,3,4\}$ (slope ratios $2.0, 3.1, 4.0$ at $r{=}2,3,4$ against the predicted $2,3,4$), where the smallest eigenvalue is rank-blind. On reduced-rank regression with closed-form $λ$, calibrated LLC recovers $λ$ at $99\%$ mean and the DDS observables rank-track it at the framework-predicted sign; on a non-linear modular-addition transformer DDS separates $d_{\mathrm{model}}$ across eighteen orders of magnitude where calibrated LLC at the protocol budget is rank-flat. Complementary to LLC's integrated posterior reading, DDS gives a directional, layer-local handle on a network's dead directions, read in closed form from its activation and gradient spectra.

奇异学习谱分析模型复杂度死方向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。