统一分析知识迁移的谱机制,揭示模型强弱背后的原理。
What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression

- 基于高维线性回归的谱分析,揭示知识迁移本质
- 知识蒸馏中扩展高频信号捕捉能力,弱到强泛化中滤除优化噪声
- 适用于理解模型压缩与泛化现象的理论框架
教师-学生知识迁移在现代机器学习中普遍存在,涵盖经典的知识蒸馏(KD)模型压缩,以及新兴的弱到强(W2S)泛化现象。现有研究虽提供零散洞见,但缺乏统一的理论框架来解释不同场景下知识迁移的有效性。本文建立高维线性回归中SGD动态的统一谱分析,阐明知识迁移在看似迥异的范式中的高效性。我们通过两种机制刻画迁移效率:在知识蒸馏中为“谱视界扩展”,使学生捕获统计上不可达的高频信号;在弱到强泛化中为“谱去噪”,学生充当优化噪声的滤波器。该框架统一了这些现象,揭示迁移效能由隐式正则化与谱域内异质学习速度之间的相互作用决定。
原文摘要 · Abstract (English)
Teacher-Student Knowledge Transfer (KT) is ubiquitous in modern machine learning, ranging from classical model compression via Knowledge Distillation (KD) to the emergent phenomenon of Weak-to-Strong (W2S) generalization. While existing studies offer isolated insights, a unified theoretical framework explaining the efficacy of KT across these disparate regimes remains lacking. In this work, we establish a unified spectral analysis of SGD dynamics in high-dimensional linear regression, elucidating the efficiency of KT across seemingly disparate regimes. We characterize KT efficiency through two distinct mechanisms: \emph{Spectral Horizon Expansion} in KD, which enables the capture of statistically inaccessible high-frequency signals, and \emph{Spectral Denoising} in W2S, where the student acts as a filter for optimization noise. Our framework unifies these phenomena, revealing that the efficacy of transfer is governed by the interplay between implicit regularization and heterogeneous spectral learning speeds over the spectrum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。