从训练数据估算非凸梯度流的群体风险曲线,提升模型泛化评估精度。
Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training Sample
- 通过近似删减路径传播删除响应,实现风险曲线估计。
- 在有限时间范围内,误差控制在 (n-1)⁻² 量级,精度高。
- 适用于宽层神经网络,无需海森矩阵可逆,适合研究泛化性能。
我们从训练样本中估计了实现的平滑非凸梯度流的条件群体风险曲线。流近似留一法(Flow-ALO)通过传播删除响应,并在近似删减路径上评估被省略的观测值。风险曲线误差分解为响应近似误差、精确留一法波动和删减到完整风险转移三部分。在固定有限时域下,有界中心化训练损失梯度、单边海森矩阵下界、局部利普希茨海森矩阵以及严格管状闭包条件,可导出删减响应误差的显式 (n−1)⁻² 界。有界评估损失梯度将删减响应界传递至得分,无需海森矩阵可逆。一阶直接杰克刀抵消与精确留一法集中控制分别处理删减到完整风险转移与波动,完成条件群体风险曲线的恢复。对于有界平滑两层均场网络且双层同时训练的情形,得分误差界在宽度上保持一致。
原文摘要 · Abstract (English)
We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample. Flow approximate leave-one-out (Flow-ALO) propagates a deletion response and evaluates omitted observations at approximate deleted paths. The risk-curve error decomposes into response approximation, exact-LOO fluctuation, and deletion-to-full risk transfer. On each fixed finite horizon, bounded centered training-loss gradients, a one-sided Hessian lower bound, locally Lipschitz Hessians, and a strict tube-closure condition yield an explicit $(n-1)^{-2}$ bound for the deletion-response error. Bounded evaluation-loss gradients transfer the deletion-response bound to the score without requiring the Hessian to be invertible. Direct first-order jackknife cancellation and exact-LOO concentration control deletion-to-full risk transfer and fluctuation, respectively, completing recovery of the conditional population-risk curve. For bounded smooth two-layer mean-field networks training both layers, the score-error bound is uniform in width.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。