arXiv:2510.21649cs.CVcs.AI2025-10

用数学曲线动态调整知识蒸馏强度,提升学生模型学习效率

A Dynamic Knowledge Distillation Method Based on the Gompertz Curve

  • 基于戈姆佩茨曲线动态调节蒸馏损失权重,匹配学生模型学习进程
  • 在CIFAR-10和CIFAR-100上分别获得最高8%和4%的准确率提升
  • 适合需要高效知识迁移的轻量化模型部署场景

本文提出一种新型动态知识蒸馏框架Gompertz-CNN,将戈姆佩茨增长模型引入训练过程,以解决传统方法难以捕捉学生模型认知能力演进的问题。针对学生模型学习初期缓慢、中期快速提升、后期饱和的特点,设计阶段感知的蒸馏策略,基于戈姆佩茨曲线动态调整蒸馏损失权重。框架融合Wasserstein距离衡量特征级差异,并通过梯度匹配对齐师生模型反向传播行为,统一于多损失目标函数中。在CIFAR-10与CIFAR-100上使用ResNet50、MobileNet_v2等不同师生架构进行大量实验,结果表明Gompertz-CNN始终优于传统蒸馏方法,在CIFAR-10上最高提升8%,CIFAR-100上最高提升4%。

原文摘要 · Abstract (English)

This paper introduces a novel dynamic knowledge distillation framework, Gompertz-CNN, which integrates the Gompertz growth model into the training process to address the limitations of traditional knowledge distillation. Conventional methods often fail to capture the evolving cognitive capacity of student models, leading to suboptimal knowledge transfer. To overcome this, we propose a stage-aware distillation strategy that dynamically adjusts the weight of distillation loss based on the Gompertz curve, reflecting the student's learning progression: slow initial growth, rapid mid-phase improvement, and late-stage saturation. Our framework incorporates Wasserstein distance to measure feature-level discrepancies and gradient matching to align backward propagation behaviors between teacher and student models. These components are unified under a multi-loss objective, where the Gompertz curve modulates the influence of distillation losses over time. Extensive experiments on CIFAR-10 and CIFAR-100 using various teacher-student architectures (e.g., ResNet50 and MobileNet_v2) demonstrate that Gompertz-CNN consistently outperforms traditional distillation methods, achieving up to 8% and 4% accuracy gains on CIFAR-10 and CIFAR-100, respectively.

知识蒸馏动态调整戈姆佩茨曲线模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。