arXiv:2608.22910stat.MLcs.LG2026-08

揭示深度网络中特征对齐的几何机制,发现其依赖传输与抵消而非单纯训练优化。

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

论文配图:A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks
图 1 · 摘自论文原文
  • 通过交换子框架量化权重协方差、门控和反向敏感度的不兼容性
  • 发现层间敏感度-协方差交换子由四种来源构成,其中负传输不平衡持续主导
  • 提出条件李雅普诺夫原理,解释风险下降不等于交换子坍缩的机制

我们构建了一个有限宽度的几何框架,描述深度神经网络中学习到的特征几何结构如何组织、传输并选择性对齐。通过三类交换子——门控与协方差之间、敏感度与协方差之间、平均梯度外积(AGOPs)与神经特征矩阵(NFMs)之间的交换子——量化了权重生成协方差、门控与反向敏感度间的不兼容性。一个精确的逐层恒等式将敏感度-协方差交换子分解为四个来源:下游传输、邻层失衡、点态敏感度波动以及非线性门控-协方差相互作用。AGOP-NFM交换子是内部交换子的奇异值加权传输,解释了为何观测到的特征侧对齐无法决定其产生的内部几何结构。缓冲局部能量解决了分离协方差子空间间的混叠问题。我们建立了谱隙、投影器演化与稳定性估计,并提出了条件李雅普诺夫原则,在显式几何误差界或内在阻尼假设下可实现衰减。这些条件并非梯度流的直接结果,阐明了风险降低未必导致交换子坍缩的原因。解析例证与数值实验展示了谱与激活几何的因子分解、瞬时增长及非零源间的抵消现象。在测试的有限时间范围内,由负传输-失衡相互作用主导的抵消在不同深度、宽度及两个回归基准上均持续存在。因此,谱对齐表现为受传输、相互作用、抵消及可能阻尼支配的层与尺度依赖现象,而非训练的普遍结果。

原文摘要 · Abstract (English)

We develop a finite-width geometric framework describing how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. Incompatibility among weight-generated covariance, gates, and backward sensitivities is quantified through three families of commutators: between gates and covariance, between sensitivities and covariance, and between average gradient outer products (AGOPs) and neural feature matrices (NFMs). An exact layerwise identity decomposes the sensitivity-covariance commutator into four sources: downstream transport, adjacent-layer imbalance, pointwise sensitivity fluctuations, and nonlinear gate-covariance interactions. The AGOP-NFM commutator is a singular-value-weighted transport of the internal commutator, explaining why observed feature-side alignment alone does not determine the internal geometry from which it emerges. Buffered localized energies resolve mixing between separated covariance subspaces. We establish spectral-gap, projector-evolution, and stabilization estimates, and formulate conditional Lyapunov principles that yield decay under explicit geometric error-bound or intrinsic-damping assumptions. These criteria do not follow from gradient flow alone and clarify why risk reduction need not imply commutator collapse. Analytic examples and numerical experiments exhibit factorization of spectral and activation geometry, transient growth, and cancellation among nonzero sources. In tested finite-time regimes, cancellation dominated by a negative transport-imbalance interaction persists across depths, widths, and two regression benchmarks. Spectral alignment therefore appears as a layer- and scale-dependent compatibility phenomenon governed by transport, interaction, cancellation, and possible damping, rather than a universal consequence of training.

深度学习几何分析交换子特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。