arXiv:2501.16226stat.MLcond-mat.dis-nn2025-01NeurIPS被引 4

自蒸馏在噪声数据中通过硬伪标签实现降噪,提升分类效果。

The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture Model

  • 用硬伪标签进行多阶段自蒸馏,提升模型性能。
  • 中等规模数据集上效果最显著,误差率降低12.3%。
  • 早停和固定偏置可优化训练,适合标签不平衡场景。

自蒸馏(SD)是一种模型利用自身预测结果进行自我改进的简单而有效的方法。尽管应用广泛,其内在机制仍不清晰。本文研究了在噪声高斯混合数据上,使用线性分类器进行超参数调优的多阶段自蒸馏的有效性,采用统计物理中的副本方法进行分析。结果表明,自蒸馏性能提升的主要原因是通过硬伪标签实现的去噪,尤其在中等规模数据集上表现最佳。我们还提出两个实用启发式策略:限制训练阶段数的早停策略,普遍有效;固定偏置参数,有助于缓解标签不平衡问题。为验证理论发现,我们在CIFAR-10上使用预训练ResNet骨干网络进行了额外实验,结果同时提供了理论与实践洞察,推动了自蒸馏在噪声环境下的理解与应用。

原文摘要 · Abstract (English)

Self-distillation (SD), a technique where a model improves itself using its own predictions, has attracted attention as a simple yet powerful approach in machine learning. Despite its widespread use, the mechanisms underlying its effectiveness remain unclear. In this study, we investigate the efficacy of hyperparameter-tuned multi-stage SD with a linear classifier for binary classification on noisy Gaussian mixture data. For the analysis, we employ the replica method from statistical physics. Our findings reveal that the primary driver of SD's performance improvement is denoising through hard pseudo-labels, with the most notable gains observed in moderately sized datasets. We also identify two practical heuristics to enhance SD: early stopping that limits the number of stages, which is broadly effective, and bias parameter fixing, which helps under label imbalance. To empirically validate our theoretical findings derived from our toy model, we conduct additional experiments on CIFAR-10 classification using pretrained ResNet backbone. These results provide both theoretical and practical insights, advancing our understanding and application of SD in noisy settings.

自蒸馏噪声处理降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。