arXiv:2605.17778math.STcs.LG2026-05

自蒸馏在特定模型中表现最优,可提升预测性能。

Self-Distillation is Optimal Among Spectral Shrinkage Estimators in Spiked Covariance Models

论文配图:Self-Distillation is Optimal Among Spectral Shrinkage Estimators in Spiked Covariance Models
图 1 · 摘自论文原文
  • 提出谱收缩估计器框架,分析自蒸馏的统计原理。
  • s步自蒸馏在含s个主成分的模型中达到最优性能。
  • 适用于理解模型优化机制,适合研究机器学习理论者。

自蒸馏已成为现代机器学习系统中提升模型性能的有前景技术。本文在带刺协方差模型中建立自蒸馏的统计基础,引入并分析一类广泛的谱收缩估计器。我们证明:对于具有s个主成分的协方差矩阵,s步自蒸馏在谱收缩估计器中达到最优性能,优于统计学和机器学习中的经典估计器。同时,我们证明s步是达到最优所必需的:任意少于s步的估计器(即(s−k)步,1≤k≤s)均严格次优。在各向同性协方差的特殊情形下,最优调参的岭回归在谱收缩估计器中表现最佳。此外,我们研究了联邦设置下多个数据中心共享谱收缩估计器,且中心服务器需聚合它们以实现最优性能的情形。在此情况下,最优本地规则仍为自蒸馏,但与数据集中于单台服务器时的最优规则不同。我们的结果揭示了自蒸馏提升预测性能的原因,并构建了其与经典收缩方法之间的广义统计联系。

原文摘要 · Abstract (English)

Self-distillation has emerged as a promising technique for improving model performance in modern machine learning systems. We develop the statistical foundations of self-distillation in spiked covariance models, by introducing and analyzing a broad class of estimators, namely spectral shrinkage estimators. We establish that for spiked covariance matrices with $s$ spikes, $s$-step self-distillation achieves optimal performance among spectral shrinkage estimators, outperforming well-known estimators in statistics and machine learning. Moreover, we show that $s$ steps are necessary for optimality: any $(s-k)$-step distilled estimator is strictly suboptimal for $1 \leq k \leq s$. For the special subclass of isotropic covariances, we show that optimally tuned Ridge regression performs best among spectral shrinkage estimators. We also study a federated approach where multiple data centers share spectral shrinkage estimators and a common server seeks to aggregate them to achieve optimal performance. In this case, we find that the best local rule again takes the form of self-distillation, though it differs from the optimal rule when data are hosted centrally on a single server. Together, our results elucidate why self-distillation improves predictive performance and provide a broader statistical framework connecting it with classical shrinkage-based methods.

自蒸馏协方差估计统计学习谱收缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。