揭示知识蒸馏中弱到强泛化机制与数据规模规律
High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws
- 在高维无正则回归下分析蒸馏过程,给出非渐近风险边界
- 弱模型蒸馏可超越强标签训练,但无法改变数据扩展规律
- 适合研究模型压缩、泛化理论与数据效率的学者参考
越来越多机器学习场景依赖知识蒸馏,即用一个代理模型的输出作为标签来监督目标模型的训练。本文针对无正则、高维回归,在两种设置下对这一过程进行精确刻画:(i) 模型偏移,代理模型任意;(ii) 分布偏移,代理模型为使用分布外数据的经验风险最小化解。在温和条件下,我们通过样本量和数据分布给出了目标模型风险的非渐近界。结果揭示了最优代理模型的形式,阐明了依数据决定性舍弃弱特征的优势与局限。在弱到强(W2S)泛化背景下,表明:(i) W2S训练(以弱模型为代理)在相同数据预算下可严格优于强标签训练;(ii) 但无法改善数据扩展规律。我们在无正则回归与神经网络架构上均验证了结果。
原文摘要 · Abstract (English)
A growing number of machine learning scenarios rely on knowledge distillation where one uses the output of a surrogate model as labels to supervise the training of a target model. In this work, we provide a sharp characterization of this process for ridgeless, high-dimensional regression, under two settings: (i) model shift, where the surrogate model is arbitrary, and (ii) distribution shift, where the surrogate model is the solution of empirical risk minimization with out-of-distribution data. In both cases, we characterize the precise risk of the target model through non-asymptotic bounds in terms of sample size and data distribution under mild conditions. As a consequence, we identify the form of the optimal surrogate model, which reveals the benefits and limitations of discarding weak features in a data-dependent fashion. In the context of weak-to-strong (W2S) generalization, this has the interpretation that (i) W2S training, with the surrogate as the weak model, can provably outperform training with strong labels under the same data budget, but (ii) it is unable to improve the data scaling law. We validate our results on numerical experiments both on ridgeless regression and on neural network architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。