用小模型指导大模型训练,显著降低复杂模型的计算成本
Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation

- 反向知识蒸馏:用小学生模型参数指导大学习者模型构建
- 在高维大数据下计算量减少70%以上,性能仍接近全样本最优
- 适用于光滑样条、核密度估计和深度学习超参优化
随着机器学习模型和数据集持续增长,构建复杂模型变得越来越计算密集。知识蒸馏通过将大型教师模型压缩为小型学生模型来降低部署成本,但无法解决教师模型构建本身是瓶颈的情形。针对这一挑战,我们提出反向知识蒸馏框架 Knowledge Cascade(KCas),利用低成本的小型学生模型信息指导更复杂教师模型的构建。尽管该方向看似反直觉(因教师通常具有更强表征能力),我们证明当有统计缩放关系支持时,学生到教师的迁移可具理论依据。首先,我们在再生核希尔伯特空间中的非参数多元函数估计中构建KCas,采用光滑样条方法,其中多平滑参数选择是主要计算瓶颈。通过渐近缩放定律,KCas将学生选定的平滑参数传递至全样本场景,大幅降低高维大规模数据的计算开销,同时保持理论保证。此外,该原则还扩展至核密度估计和深度学习超参数转移。模拟与真实数据实验表明,KCas实现显著计算节省,性能强,甚至有时优于对应全样本方法。
原文摘要 · Abstract (English)
As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge distillation reduces deployment cost by compressing a large, well-trained teacher model into a compact student model, but it does not address settings where constructing the teacher itself is the bottleneck. Motivated by this challenge, we introduce Knowledge Cascade (KCas), a reverse knowledge distillation framework that uses information from a small, inexpensive student model to guide the development of a more complex teacher model. Although this direction is counterintuitive because the teacher typically has greater representational capacity, we show that student-to-teacher transfer can be principled when supported by statistical scaling relationships. We first develop KCas for nonparametric multivariate functional estimation in reproducing kernel Hilbert spaces via smoothing splines, where selecting multiple smoothing parameters is a major computational bottleneck. KCas transfers student-selected smoothing parameters to the full-sample regime through asymptotic scaling laws, substantially reducing computational cost for high-dimensional and large-scale datasets while retaining theoretical guarantees. Beyond smoothing splines, we illustrate the same principle through kernel density estimation and deep learning hyperparameter transfer. Simulations and real-data experiments show that KCas achieves substantial computational savings while maintaining strong statistical performance, and can sometimes outperform the corresponding full-sample procedure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。