提出自正则化框架,统一解释梯度下降等算法的隐式复杂度控制机制。
Self-Regularized Learning Methods
- 基于自正则化概念,无需显式正则项即可控制模型复杂度
- 证明算法达到极小极大最优泛化率,只需满足自正则化条件
- 适用于数据驱动超参数选择,支持核方法中的早停策略
我们提出一个基于自正则化概念的通用学习算法分析框架,捕捉无需显式正则化即可实现的隐式复杂度控制。该框架源于对梯度下降等算法隐式正则化现象的观察:自正则化算法中,预测器的复杂度由达成相同经验风险的最简比较器决定。该框架涵盖经典正则化经验风险最小化与梯度下降。在此基础上,我们对这类算法进行了全面的统计分析,证明其可达到极小极大最优泛化率——仅需验证算法具备自正则化特性,其余条件均由学习问题自身决定。最后,我们讨论了数据依赖的超参数选择问题,给出一个通用结果,可在双对数因子内实现极小极大最优率,并覆盖基于RKHS的梯度下降中的数据驱动早停策略。
原文摘要 · Abstract (English)
We introduce a general framework for analyzing learning algorithms based on the notion of self-regularization, which captures implicit complexity control without requiring explicit regularization. This is motivated by previous observations that many algorithms, such as gradient-descent based learning, exhibit implicit regularization. In a nutshell, for a self-regularized algorithm the complexity of the predictor is inherently controlled by that of the simplest comparator achieving the same empirical risk. This framework is sufficiently rich to cover both classical regularized empirical risk minimization and gradient descent. Building on self-regularization, we provide a thorough statistical analysis of such algorithms including minmax-optimal rates, where it suffices to show that the algorithm is self-regularized -- all further requirements stem from the learning problem itself. Finally, we discuss the problem of data-dependent hyperparameter selection, providing a general result which yields minmax-optimal rates up to a double logarithmic factor and covers data-driven early stopping for RKHS-based gradient descent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。