动态选择导师并调整教学策略,提升知识蒸馏效果
Classroom-Inspired Multi-Mentor Distillation with Adaptive Learning Strategies
- 根据样本表现动态筛选高质导师,避免错误累积
- 按学生与导师差距调节教学力度,控制学习节奏
- 在图像分类与人体姿态估计任务中均超越现有方法
我们提出 ClassroomKD,一种受课堂环境启发的多导师知识蒸馏框架,用于提升学生模型与多个不同知识水平导师之间的知识迁移效率。不同于依赖固定师生关系的传统方法,该框架根据每个数据样本的表现,动态选择并自适应调整不同导师的教学策略。框架包含两个核心模块:知识过滤(KF)模块和指导模块。KF 模块基于每个输入样本的性能动态排序导师,仅激活高质量导师以减少误差积累并防止信息丢失;指导模块通过调节各导师对学生的影响力,依据学生与导师间的动态性能差距,有效调控学习速率。在图像分类(CIFAR-100、ImageNet)和2D人体姿态估计(COCO Keypoints、MPII Human Pose)任务上的大量实验表明,ClassroomKD 在多种网络架构下均优于现有知识蒸馏方法。结果表明,动态且自适应的导师选择与引导机制可实现更高效的知识迁移,为通过蒸馏提升模型性能提供新路径。
原文摘要 · Abstract (English)
We propose ClassroomKD, a novel multi-mentor knowledge distillation framework inspired by classroom environments to enhance knowledge transfer between the student and multiple mentors with different knowledge levels. Unlike traditional methods that rely on fixed mentor-student relationships, our framework dynamically selects and adapts the teaching strategies of diverse mentors based on their effectiveness for each data sample. ClassroomKD comprises two main modules: the Knowledge Filtering (KF) module and the Mentoring module. The KF Module dynamically ranks mentors based on their performance for each input, activating only high-quality mentors to minimize error accumulation and prevent information loss. The Mentoring Module adjusts the distillation strategy by tuning each mentor's influence according to the dynamic performance gap between the student and mentors, effectively modulating the learning pace. Extensive experiments on image classification (CIFAR-100 and ImageNet) and 2D human pose estimation (COCO Keypoints and MPII Human Pose) demonstrate that ClassroomKD outperforms existing knowledge distillation methods for different network architectures. Our results highlight that a dynamic and adaptive approach to mentor selection and guidance leads to more effective knowledge transfer, paving the way for enhanced model performance through distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。