arXiv:2503.11363cs.SDcs.LG2025-03被引 1

探究教师模型特性对声学场景分类压缩效果的影响

Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification

  • 对比不同教师网络结构、大小及训练方法
  • 发现教师规模与设备泛化策略显著影响学生性能
  • 适合关注模型压缩与知识蒸馏的开发者

知识蒸馏(KD)是将大型模型知识压缩为更小高效模型的常用技术,在构建高性能低复杂度声学场景分类(ASC)系统中表现优异,近三年所有DCASE挑战赛顶级提交均采用此方法。尽管已有大量研究聚焦于蒸馏流程设计、学生模型优化及教师集成方法,但关于教师模型哪些属性对低复杂度学生更有利的研究仍不足。本文通过实验分析不同教师网络架构、教师模型大小、设备泛化训练方法以及集成策略对学生性能的影响。结果表明,教师模型大小、设备泛化方法、集成策略及集成规模是决定学生网络性能的关键因素。

原文摘要 · Abstract (English)

Knowledge Distillation (KD) is a widespread technique for compressing the knowledge of large models into more compact and efficient models. KD has proved to be highly effective in building well-performing low-complexity Acoustic Scene Classification (ASC) systems and was used in all the top-ranked submissions to this task of the annual DCASE challenge in the past three years. There is extensive research available on establishing the KD process, designing efficient student models, and forming well-performing teacher ensembles. However, less research has been conducted on investigating which teacher model attributes are beneficial for low-complexity students. In this work, we try to close this gap by studying the effects on the student's performance when using different teacher network architectures, varying the teacher model size, training them with different device generalization methods, and applying different ensembling strategies. The results show that teacher model sizes, device generalization methods, the ensembling strategy and the ensemble size are key factors for a well-performing student network.

知识蒸馏声学分类模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。