ZeNN通过新架构突破传统MLP在宽网络下的局限,提升高频特征学习能力。
A ZeNN architecture to avoid the Gaussian trap
- 基于调和分析设计新架构,引入可枚举感知机与非学习权重以保障收敛
- 无限宽时实现逐点收敛,突破高斯性限制,保留非高斯结构与特征学习能力
- 有限宽度下仍能高效学习低维函数的高频特征,适合高频率建模任务
我们提出一种新型简单架构——泽塔神经网络(ZeNN),以克服标准多层感知机(MLP)的若干缺陷。在宽网络极限下,传统MLP表现为非参数化,缺乏明确定点极限,丧失非高斯特性并无法进行特征学习;且有限宽度的MLP在学习高频特征方面表现不佳。新架构受调和分析三个原则启发:(i)枚举感知机并引入非学习权重以保证收敛;(ii)引入缩放(或频率)因子;(iii)选择能生成近正交系统的激活函数。我们证明这些设计可解决上述问题:在无限宽极限下,ZeNN实现逐点收敛,展现出超越高斯性的丰富渐近结构,并具备特征学习能力;当选择合适激活函数时,有限宽度的ZeNN在低维域函数的高频特征学习上表现优异。
原文摘要 · Abstract (English)
We propose a new simple architecture, Zeta Neural Networks (ZeNNs), in order to overcome several shortcomings of standard multi-layer perceptrons (MLPs). Namely, in the large width limit, MLPs are non-parametric, they do not have a well-defined pointwise limit, they lose non-Gaussian attributes and become unable to perform feature learning; moreover, finite width MLPs perform poorly in learning high frequencies. The new ZeNN architecture is inspired by three simple principles from harmonic analysis: i) Enumerate the perceptons and introduce a non-learnable weight to enforce convergence; ii) Introduce a scaling (or frequency) factor; iii) Choose activation functions that lead to near orthogonal systems. We will show that these ideas allow us to fix the referred shortcomings of MLPs. In fact, in the infinite width limit, ZeNNs converge pointwise, they exhibit a rich asymptotic structure beyond Gaussianity, and perform feature learning. Moreover, when appropriate activation functions are chosen, (finite width) ZeNNs excel at learning high-frequency features of functions with low dimensional domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。