找出了线性回归中最佳隐式偏置,提升过参数化模型泛化能力。
Optimal Implicit Bias in Linear Regression
- 通过凸函数最小化分析过参数线性回归的隐式偏置机制
- 给出在特定条件下最优泛化误差的紧下界
- 适用于研究泛化性能与优化算法关系的研究者
现代学习问题多为过参数化,参数量远超训练样本数。在此情形下,训练损失存在无穷多个全局最优解,均完全拟合数据但泛化性能各异,最终收敛到哪个解取决于优化算法的隐式偏置。本文探讨的问题是:何种隐式偏置能带来最优泛化性能?为此,我们对非各向同性高斯数据下的过参数化线性回归中凸函数/势函数最小化所得插值器的泛化性能进行了精确渐近分析。特别地,我们以过参数化比、标签噪声方差、数据协方差的特征谱和待估计参数的先验分布为参数,推导出该类插值器中可能达到的最优泛化误差的紧下界。最后,在涉及高斯卷积先验的对数凹性满足的充分条件下,找到了可实现该下界的最优凸隐式偏置。
原文摘要 · Abstract (English)
Most modern learning problems are over-parameterized, where the number of learnable parameters is much greater than the number of training data points. In this over-parameterized regime, the training loss typically has infinitely many global optima that completely interpolate the data with varying generalization performance. The particular global optimum we converge to depends on the implicit bias of the optimization algorithm. The question we address in this paper is, ``What is the implicit bias that leads to the best generalization performance?". To find the optimal implicit bias, we provide a precise asymptotic analysis of the generalization performance of interpolators obtained from the minimization of convex functions/potentials for over-parameterized linear regression with non-isotropic Gaussian data. In particular, we obtain a tight lower bound on the best generalization error possible among this class of interpolators in terms of the over-parameterization ratio, the variance of the noise in the labels, the eigenspectrum of the data covariance, and the underlying distribution of the parameter to be estimated. Finally, we find the optimal convex implicit bias that achieves this lower bound under certain sufficient conditions involving the log-concavity of the distribution of a Gaussian convolved with the prior of the true underlying parameter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。