用几何框架统一解释深度网络权重对齐现象,揭示其内在数学结构。
Flag Varieties: A Geometric Framework for Deep Network Alignment
- 基于几何不变理论,证明权重对齐由旗流形决定,子空间交维是唯一不变量。
- 权重衰减使对齐以指数速率加速,非线性激活产生阻碍精确对齐的反对易项。
- 无需前向传播即可通过权重空间观测内部对齐结构,适合模型诊断与分析。
对齐,即深层网络中相邻权重矩阵倾向于形成兼容的子空间方向,是梯度流动、神经坍缩及跨架构表征相似性的基础。尽管已有大量实证研究,但这些现象缺乏统一的理论解释——现有理论多为事后拟合,针对特定观察使用不同数学工具。本文反其道而行之,推导出层间对齐所必然要求的数学结构。利用几何不变理论,我们证明对齐几何具有一个标准的闭合、多稳定分支,由旗流形(flag variety)给出,且子空间交维是其唯一的参数化无关可观测量,确立了子空间度量并非经验惯例而是数学必然。该统一框架带来两个动力学后果:岭正则化以权重衰减设定的指数速率驱动子空间对齐;而非线性激活引入反对易项阻碍精确基对齐,这在非线性网络中普遍存在,而在线性网络中不存在。二者共同从第一原理解释了神经坍缩中的层级2/3结构,而非事后分析。反对易项大小与头子空间重叠进一步作为权重空间中的内部对齐窗口,无需前向计算。在多层感知机、残差网络及预训练语言模型上的实验验证了所提诊断方法的有效性,并明确了其适用范围。
原文摘要 · Abstract (English)
Alignment, the tendency of adjacent weight matrices in deep networks to develop compatible subspace orientations, underlies gradient flow, Neural Collapse, and representation similarity across architectures. Despite extensive empirical documentation, these phenomena have resisted unified theoretical treatment: existing explanations are post-hoc, each fitted to a specific observation with whatever mathematics is at hand. We reverse this direction by deriving the mathematical structure that layerwise alignment inherently demands. Using geometric invariant theory, we prove that alignment geometry has a canonical closed, polystable stratum given by a flag variety, and that subspace intersection dimension is its unique reparameterization-invariant observable, establishing that subspace metrics are not empirical conventions but mathematical necessities. This unified framework yields two dynamical consequences: ridge regularization drives subspace alignment at an exponential rate set by weight decay, whereas nonlinear activations induce a commutator obstruction to exact basis alignment, generically present in nonlinear networks and absent in linear ones. Together these give a geometric explanation of the Level-2/3 hierarchy in Neural Collapse from first principles rather than post-hoc analysis. The commutator magnitude and head subspace overlap further serve as weight-space windows into internal alignment structure, requiring no forward passes. Experiments on multilayer perceptrons, residual networks, and pretrained language models support the proposed diagnostics and delineate their scope.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。