arXiv:2409.17312cs.CLcs.LG2024-09被引 8

小模型通过集成蒸馏,数据少也能超越大模型。

BabyLlama-2: Ensemble-Distilled Models Consistently Outperform Teachers With Limited Data

  • 用两个教师模型蒸馏出3450万参数的小模型。
  • 在1000万词数据下表现超过1000万词训练的基线和教师模型。
  • 证明蒸馏优势非因教师超参不佳,适合小数据场景研究者参考。

我们提出BabyLlama-2,一个3.45亿参数的模型,通过在1000万词语料上对两个教师模型进行集成蒸馏,用于BabyLM竞赛。在BLiMP和SuperGLUE基准测试中,BabyLlama-2的表现优于使用相同数据混合的1000万和1000万词数据集训练的基线模型,也优于其教师模型。通过大规模超参数搜索,我们证明蒸馏的优势并非源于教师模型的次优超参数选择。研究结果强调需进一步探索蒸馏技术,尤其是在数据受限的场景下。

原文摘要 · Abstract (English)

We present BabyLlama-2, a 345 million parameter model distillation-pretrained from two teachers on a 10 million word corpus for the BabyLM competition. On BLiMP and SuperGLUE benchmarks, BabyLlama-2 outperforms baselines trained on both 10 and 100 million word datasets with the same data mix, as well as its teacher models. Through an extensive hyperparameter sweep, we demonstrate that the advantages of distillation cannot be attributed to suboptimal hyperparameter selection of the teachers. Our findings underscore the need for further investigation into distillation techniques, particularly in data-limited settings.

模型蒸馏小样本学习语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。