量化了正则化路径下分类模型的过拟合程度,揭示参数λ如何影响泛化误差。
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
- 基于任意先验和描述语言,完整刻画了修正版两部分编码最小描述长度的正则化曲线。
- 精确给出最坏情况下的极限误差随正则化参数λ和噪声水平的变化关系。
- 揭示了λ=1时欠正则化的严重性,适合关注泛化性能与正则化设计的研究者。
我们基于任意先验或描述语言,对二分类任务中一种修正的两部分编码最小描述长度(MDL)学习规则的完整正则化曲线进行了全面表征。Grunwald 和 Langford [2004] 从前置的统计学习理论(频繁统计最坏情形)角度指出,当惩罚参数λ=1时,该MDL规则缺乏渐近一致性,表明其存在欠正则化问题。为深入理解欠正则化和过拟合的良性和灾难性程度,本文精确量化了最坏情况下极限误差随正则化参数λ和噪声水平(或近似误差)的变化关系,显著收紧了Grunwald和Langford对λ=1情形的分析,并将其推广至所有λ取值。
原文摘要 · Abstract (English)
We provide a complete characterization of the entire regularization curve of a modified two-part-code Minimum Description Length (MDL) learning rule for binary classification, based on an arbitrary prior or description language. Grunwald and Langford [2004] previously established the lack of asymptotic consistency, from an agnostic PAC (frequentist worst case) perspective, of the MDL rule with a penalty parameter of $λ=1$, suggesting that it underegularizes. Driven by interest in understanding how benign or catastrophic under-regularization and overfitting might be, we obtain a precise quantitative description of the worst case limiting error as a function of the regularization parameter $λ$ and noise level (or approximation error), significantly tightening the analysis of Grunwald and Langford for $λ=1$ and extending it to all other choices of $λ$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。