arXiv:2412.08347cs.CLcs.AI2024-12被引 3

调高学习率与批次比能提升小模型推理能力

SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs

  • 通过调整学习率与批次比,优化小模型对推理任务的适应性
  • 在GSM8K上达51.6%准确率,比同类模型高出3.4%
  • 适合关注小模型高效对齐与推理增强的研究者

我们提出SmolTulu-1.7b-Instruct(本报告中称SmolTulu-DPO-1130),基于AllenAI的Tulu 3后训练流程,适配HuggingFace的SmolLM2-1.7B基础模型。通过135M参数模型的全面实证分析,发现学习率与批次大小之比对模型性能有显著影响,且任务依赖性强:推理任务如ARC和GSM8K在较高比率下表现更优,而模式识别任务如HellaSwag和IFEval则在较低比率下达到最佳。该发现指导了SmolTulu的开发,在指令遵循任务上取得亚20亿参数模型中的领先水平,IFEval得分67.7%(Δ11%),数学推理达GSM8K 51.6%(Δ3.4%),另一版本在ARC上达57.1%(Δ5.4%)。我们公开模型、训练方案及消融研究,推动高效模型对齐的进一步探索,证明合理调整优化动态可缩小小模型与大模型的能力差距。

原文摘要 · Abstract (English)

We present SmolTulu-1.7b-Instruct, referenced in this report as SmolTulu-DPO-1130, an instruction-tuned language model that adapts AllenAI's Tulu 3 post-training pipeline to enhance Huggingface's SmolLM2-1.7B base model. Through comprehensive empirical analysis using a 135M parameter model, we demonstrate that the relationship between learning rate and batch size significantly impacts model performance in a task-dependent manner. Our findings reveal a clear split: reasoning tasks like ARC and GSM8K benefit from higher learning rate to batch size ratios, while pattern recognition tasks such as HellaSwag and IFEval show optimal performance with lower ratios. These insights informed the development of SmolTulu, which achieves state-of-the-art performance among sub-2B parameter models on instruction following, scoring 67.7% on IFEval ($Δ$11%), and mathematical reasoning with 51.6% on GSM8K ($Δ$3.4%), with an alternate version achieving scoring 57.1% on ARC ($\Delta5.4%$). We release our model, training recipes, and ablation studies to facilitate further research in efficient model alignment, demonstrating that careful adaptation of optimization dynamics can help bridge the capability gap between small and large language models.

小模型推理增强优化策略指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。