用反KL散度让小模型专注学习,单教师效果胜过多个教师。
Choosy Babies Need One Coach: Inducing Mode-Seeking Behavior in BabyLlama with Reverse KL Divergence
- 改用反KL散度引导学生模型聚焦特定模式,避免平均化。
- 单教师设置在多数任务上表现优于或多于双教师模型。
- 结合优化策略后模型性能提升,适合资源受限场景的高效训练。
本研究提交至第二届BabyLM挑战赛的严格小型模型赛道。采用以BabyLLaMa模型(Timiryasov and Tastet, 2023)为骨干的师生蒸馏框架。为使学生学习更聚焦,将目标函数替换为反Kullback-Leibler散度,该方法可促使计算学习者产生模式聚焦行为(而非模式平均)。我们进一步实验仅使用单个教师(而非两个教师的集成),并引入额外优化策略以提升蒸馏效果。实验表明,在反KL散度下,单教师模型在多数任务中表现优于或匹配多教师模型。同时,结合先进优化技术可进一步提升模型性能,验证了所提方法的有效性与鲁棒性。这些发现支持我们的观点:‘挑食的小孩只需一个导师’。
原文摘要 · Abstract (English)
This study presents our submission to the Strict-Small Track of the 2nd BabyLM Challenge. We use a teacher-student distillation setup with the BabyLLaMa model (Timiryasov and Tastet, 2023) as a backbone. To make the student's learning process more focused, we replace the objective function with a reverse Kullback-Leibler divergence, known to cause mode-seeking (rather than mode-averaging) behaviour in computational learners. We further experiment with having a single teacher (instead of an ensemble of two teachers) and implement additional optimization strategies to improve the distillation process. Our experiments show that under reverse KL divergence, a single-teacher model often outperforms or matches multiple-teacher models across most tasks. Additionally, incorporating advanced optimization techniques further enhances model performance, demonstrating the effectiveness and robustness of our proposed approach. These findings support our idea that "choosy babies need one coach".
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。