研究浅层网络镜像流的隐式偏置,揭示其与梯度流的相似性及可调控性。
Implicit Bias of Mirror Flow for Shallow Neural Networks in Univariate Regression
- 通过变分问题刻画ReLU网络镜像流的隐式偏置机制。
- 在无限宽条件下,镜像流呈现懒惰训练,且偏置与普通梯度流一致。
- 引入缩放势函数,可实现非核范数的丰富偏置,适合自适应曲率惩罚场景。
我们研究了在单变量最小二乘回归中,宽而浅的神经网络下镜像流的隐式偏置。对于一大类势函数,当网络宽度趋于无穷时,镜像流表现出懒惰训练,其隐式偏置与普通梯度流相同。对于ReLU网络,我们通过函数空间中的变分问题刻画了这一偏置。分析涵盖普通梯度流的先前结果,并克服了需不可行数据调整或跳跃连接的局限。我们进一步引入缩放势函数,发现此时镜像流仍呈现懒惰训练但不在核范数范畴。对于绝对值激活网络,缩放势函数诱导出一类通常无法由RKHS范数描述的丰富偏置。核心启示是:参数初始化决定不同输入位置上曲率惩罚强度,而缩放势函数决定不同曲率大小的惩罚方式。
原文摘要 · Abstract (English)
We examine the implicit bias of mirror flow in univariate least squares error regression with wide and shallow neural networks. For a broad class of potential functions, we show that mirror flow exhibits lazy training and has the same implicit bias as ordinary gradient flow when the network width tends to infinity. For ReLU networks, we characterize this bias through a variational problem in function space. Our analysis includes prior results for ordinary gradient flow as a special case and lifts limitations which required either an intractable adjustment of the training data or networks with skip connections. We further introduce scaled potentials and show that for these, mirror flow still exhibits lazy training but is not in the kernel regime. For networks with absolute value activations, we show that mirror flow with scaled potentials induces a rich class of biases, which generally cannot be captured by an RKHS norm. A takeaway is that whereas the parameter initialization determines how strongly the curvature of the learned function is penalized at different locations of the input space, the scaled potential determines how the different magnitudes of the curvature are penalized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。