用随机效应模型重新解释梯度流训练,实现可推断的早停与训练时长决策。
Gradient-Flow Optimization as Dynamic Random-Effects Inference: Testing and Early Stopping with Applications to Deep Learning

- 将梯度流视为随机效应模型下的最优线性预测,训练时间变更为方差分量参数。
- REML引导的早停规则使优化谱损失与训练算子特征值解相关,提升预测精度。
- 适用于固定核深度学习场景,减少对验证集和多次检查点的依赖。
梯度流优化通常被视为最小化经验损失的算法过程,训练时长通过验证或启发式早停规则选择。我们为梯度流训练构建了统计推断框架。当拟合值通过一个时不变的半正定训练算子演化时,每个时刻的输出等价于对应随机效应模型下的最佳线性无偏预测。训练时间成为调控残差噪声与结构信号间方差分配的方差分量参数。这将两个训练决策转化为推断问题:是否需要训练转化为初始值之外是否存在信号的方差分量检验;训练时长则转化为训练时间方差分量的受限最大似然(REML)估计。我们证明,由REML指导的早停规则在优化谱损失与训练算子特征值解相关时达到最优。该规则在固定设计的样本内风险和随机设计的样本外风险下具有渐近预测最优性。固定核梯度范式下的深度学习模型是本结果的典型实例。数值实验及英国生物银行蛋白质组学应用显示,基于REML的早停方法在降低对验证集依赖和重复检查点评估的同时,仍保持竞争力的准确率。
原文摘要 · Abstract (English)
Gradient-flow optimization is usually viewed as an algorithmic procedure for minimizing empirical loss, with training duration selected by validation or heuristic early stopping rules. We develop a statistical inference framework for gradient-flow training. We show that whenever fitted values evolve through a time-invariant positive semidefinite training operator, the output at each time is equivalent to the best linear unbiased predictor under a corresponding random-effects model. Training time then becomes a variance-component parameter governing variance reallocation from residual noise to structured signal. This turns two training decisions into inferential problems: whether training is needed becomes a variance-component test for signal beyond initialization, and how long to train becomes restricted maximum likelihood (REML) estimation of the training-time variance component. We show that the REML-guided early stopping rule selects the time at which optimized spectral losses become decorrelated from the training-operator eigenvalues. The asymptotic prediction optimality of the REML-guided early stopping time is established for fixed-design in-sample risk and random-design out-of-sample risk. Deep learning models in fixed-kernel gradient regimes provide canonical instantiations for our results. Numerical experiments and a UK Biobank proteomics application show competitive accuracy of the REML-guided early stopping time with reduced reliance on validation splits and repeated checkpoint evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。