发现深度ReLU网络多数参数可被唯一确定,突破传统对称性限制。
Most ReLU Networks Admit Identifiable Parameters
- 基于加权多面体复形分析网络参数冗余
- 宽度≥2的层在开集上参数可识别,功能维度等于参数数减隐层神经元数
- 即使最优表示仍有非平凡冗余,且深层网络有通用深度优势
我们研究深度ReLU网络的函数实现映射,关注何时函数能唯一确定其参数(忽略缩放与置换)。为分析超出标准对称性的隐藏冗余,引入基于加权多面体复形的框架。主要结果表明:对于输入和隐藏层宽度均至少为2的任意架构,存在一个开集上的可识别参数。这意味着此类架构的功能维度恰好等于参数数减去隐藏神经元数。进一步证明,即使是最小功能表示仍可能具有非平凡参数冗余。最后,建立了一般性深度层次结构:对某个开集参数,所实现函数无法由任何更浅网络通用表示。
原文摘要 · Abstract (English)
We study the realization map of deep ReLU networks, focusing on when a function determines its parameters up to scaling and permutation. To analyze hidden redundancies beyond these standard symmetries, we introduce a framework based on weighted polyhedral complexes. Our main result shows that for every architecture whose input and hidden layers have width at least two, there exists an open set of identifiable parameters. This implies that the functional dimension of every such architecture is exactly the number of parameters minus the number of hidden neurons. We further show that minimal functional representations can still have non-trivial parameter redundancies. Finally, we establish a generic depth hierarchy, whereby for an open set of parameters the realized function cannot be represented generically by any shallower network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。