arXiv:2601.21645cs.LGmath.CT2026-01中稿 · ICML被引 7

证明了可识别的等变网络各层也等变,解释了训练中等变结构为何自然出现。

Identifiable Equivariant Networks are Layerwise Equivariant

  • 通过可识别性假设,将整体等变性转化为逐层等变性。
  • 在典型网络架构中,层间等变结构可被数学保证。
  • 为深度网络训练中等变权重的涌现提供理论依据。

我们研究了深度神经网络中端到端等变性与逐层等变性之间的关系。证明:若一个网络的端到端函数对输入和输出空间上的群作用保持等变,那么存在一种参数选择,使得该网络的各层在潜在空间上对某些群作用也保持等变,且整体函数不变。该结果依赖于模型参数在适当意义下的可识别性,这一性质已在大量网络架构中被证明成立,而对其他架构仍为猜想。我们的理论基于抽象形式化框架,不依赖具体网络结构。总体而言,本研究为训练过程中神经网络权重中等变结构的自然出现提供了数学解释——这一现象在实践中反复被观察到。

原文摘要 · Abstract (English)

We investigate the relation between end-to-end equivariance and layerwise equivariance in deep neural networks. We prove the following: For a network whose end-to-end function is equivariant with respect to group actions on the input and output spaces, there is a parameter choice yielding the same end-to-end function such that its layers are equivariant with respect to some group actions on the latent spaces. Our result assumes that the parameters of the model are identifiable in an appropriate sense. This identifiability property has been established in the literature for a large class of networks, to which our results apply immediately, while it is conjectural for others. The theory we develop is grounded in an abstract formalism, and is therefore architecture-agnostic. Overall, our results provide a mathematical explanation for the emergence of equivariant structures in the weights of neural networks during training -- a phenomenon that is consistently observed in practice.

等变网络可识别性理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。