arXiv:2510.15516cs.LGcs.AI2025-10被引 1

发现知识蒸馏在小数据下效果更显著,挑战了传统理解。

Revisiting Knowledge Distillation: The Hidden Role of Dataset Size

  • 研究发现蒸馏在小数据集上效果更强,提出'数据效率'新特性。
  • 实验证明蒸馏并非简单标签平滑,支持'暗知识'理论。
  • 适合关注模型压缩与小样本学习的研究者参考。

知识蒸馏(KD)是深度学习中通过教师模型指导学生模型训练的常用技术,但其工作机制仍不明确。现有研究主要关注模型规模和泛化能力,本文首次从数据集规模这一新维度切入,通过跨多种数据集、任务和神经网络架构的实验,发现蒸馏效果不仅在低数据场景中得以保留,反而被放大,提出“数据效率”这一新特性。基于此,我们检验了现有蒸馏理论的预测能力,结果否定了蒸馏等同于标签平滑的假设,并进一步支持了‘暗知识’假说。此外,分析了目标函数、模型规模及样本相对数量等建模因素的影响。研究表明,数据集大小可能是理解蒸馏机制时被忽视的关键变量。

原文摘要 · Abstract (English)

The concept of knowledge distillation (KD) describes the training of a student model from a teacher model and is a widely adopted technique in deep learning. However, it is still not clear how and why distillation works. Previous studies focus on two central aspects of distillation: model size, and generalisation. In this work we study distillation in a third dimension: dataset size. We present a suite of experiments across a wide range of datasets, tasks and neural architectures, demonstrating that the effect of distillation is not only preserved but amplified in low-data regimes. We call this newly discovered property the data efficiency of distillation. Equipped with this new perspective, we test the predictive power of existing theories of KD as we vary the dataset size. Our results disprove the hypothesis that distillation can be understood as label smoothing, and provide further evidence in support of the dark knowledge hypothesis. Finally, we analyse the impact of modelling factors such as the objective, scale and relative number of samples on the observed phenomenon. Ultimately, this work reveals that the dataset size may be a fundamental but overlooked variable in the mechanisms underpinning distillation.

知识蒸馏小样本学习模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。