揭示神经层扩散中过度平滑的本质是表示退化,提出用几何不变量理论改进模型稳定性。
Oversmoothing as Representation Degeneracy in Neural Sheaf Diffusion
- 将图上的层扩散视为射影表示,用组合代数视角分析其表示结构。
- 发现当层维数相等时,学习到的几何会退化为平凡分量,导致信息丢失。
- 引入基于矩映射的正则项,打破对称性可提升模型稳定性和泛化性能。
神经层扩散(NSD)通过将标量图拉普拉斯矩阵替换为可学习的层拉普拉斯矩阵,扩展了基于扩散的图神经网络,其学习的限制映射定义了任务自适应的几何结构。尽管已知NSD的扩散极限是全局截面空间,但该调和空间的表示论结构仍不明确。本文通过将图上的细胞层与关联的关联有向图表示相对应,建立了一种范畴论解释。在此对应下,学习到的层几何成为有限维表示空间中的点。我们证明,底层关联有向图表示的直和分解会诱导出扩散极限中调和空间的分解。这给出了过度平滑作为表示退化的代数解释:学习到的层可能坍缩至低复杂度的直和分量,导致全局截面无法保留判别信息。基于此观点,我们将层扩散与几何不变量理论中的稳定性与矩映射原理相联系,提出受矩映射启发的正则项,引导限制映射趋向平衡的表示几何。我们识别出在等维层架构中存在结构性障碍:当节点和边的层维数 $d_v = d_e$ 时,学习到的稳定性参数的容许性迫使平凡全对象分量位于稳定性壁上。非均匀层维数可消除此障碍,使自适应稳定性具有意义。在异质性基准测试上的实验结果与此机制一致:打破层对称性可降低方差或改善验证表现,自适应稳定性在选定矩形设置中更有效。总体而言,我们的框架将过度平滑重新诠释为学习层扩散中表示几何的退化现象。
原文摘要 · Abstract (English)
Neural Sheaf Diffusion (NSD) generalizes diffusion-based Graph Neural Networks by replacing scalar graph Laplacians with sheaf Laplacians whose learned restriction maps define a task-adapted geometry. While the diffusion limit of NSD is known to be the space of global sections, the representation-theoretic structure of this harmonic space remains largely implicit. We develop a quiver-theoretic interpretation of NSD by identifying cellular sheaves on graphs with representations of the associated incidence quiver. Under this correspondence, learned sheaf geometries become points in a finite-dimensional representation space. We show that direct-sum decompositions of the underlying incidence-quiver representation induce decompositions of the harmonic space reached in the diffusion limit. This gives an algebraic interpretation of oversmoothing as representation degeneration: learned sheaves may collapse toward low-complexity summands whose global sections fail to preserve discriminative information. Building on this viewpoint, we connect sheaf diffusion to stability and moment-map principles from Geometric Invariant Theory. We introduce moment-map-inspired regularizers that bias restriction maps toward balanced representation geometries, and identify a structural obstruction in equal-stalk architectures: when $d_v = d_e$, admissibility for learnable stability parameters forces the trivial all-object summand onto a stability wall. Non-uniform stalk dimensions remove this obstruction, making adaptive stability meaningful. Experiments on heterophilic benchmarks are consistent with this mechanism: breaking stalk symmetry can reduce variance or improve validation behavior, and adaptive stability becomes more effective in selected rectangular settings. Overall, our framework reframes oversmoothing as a degeneration phenomenon in the representation geometry underlying learned sheaf diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。