通过分布平衡提升无标签多教师知识蒸馏效果
PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation
- 引入PHI-S标准化方法,统一多教师激活分布尺度
- 在多个数据集上使学生模型性能优于现有方法
- 适合研究无监督知识蒸馏与多教师协同的学者
多种视觉基础模型具有不同的优势与缺陷,可通过无标签的异构多教师知识蒸馏进行改进,这类方法称为“聚集模型”。本文研究教师激活统计特性的影响,特别是损失函数对最终学生模型质量的作用。探索标准的统计归一化技术以更好对齐不同分布,并评估其效果。进一步分析下游教师匹配指标的影响,从而提出使用Hadamard矩阵。利用这些矩阵,证明其具备各向同性标准化特性,即对多维分布的每一维度采用相同尺度进行标准化。该方法称为“PHI标准化”(PHI-S),实证表明其在所研究的方法中表现最优。
原文摘要 · Abstract (English)
Various visual foundation models have distinct strengths and weaknesses, both of which can be improved through heterogeneous multi-teacher knowledge distillation without labels, termed "agglomerative models." We build upon this body of work by studying the effect of the teachers' activation statistics, particularly the impact of the loss function on the resulting student model quality. We explore a standard toolkit of statistical normalization techniques to better align the different distributions and assess their effects. Further, we examine the impact on downstream teacher-matching metrics, which motivates the use of Hadamard matrices. With these matrices, we demonstrate useful properties, showing how they can be used for isotropic standardization, where each dimension of a multivariate distribution is standardized using the same scale. We call this technique "PHI Standardization" (PHI-S) and empirically demonstrate that it produces the best student model across the suite of methods studied.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。