通过弱化指令微调,提升模型少样本学习能力。
Improving Instruct Models for Free: A Study on Partial Adaptation
- 采用部分适配方法降低指令微调强度。
- 在多种模型上显著提升少样本上下文学习表现。
- 适合关注上下文学习性能的实践者参考。
指令模型通常比基础模型更优,具备更强的指令遵循能力。然而,指令微调可能导致预训练知识遗忘,或使模型过度对话化、冗长,从而降低上下文少样本学习性能。本文通过部分适配方法减弱指令微调强度,研究基线与指令模型之间的性能演变。结果表明,在多个模型家族和尺寸下,降低微调强度可显著提升覆盖多种经典自然语言任务的少样本上下文学习基准表现。代价是指令遵循能力略有下降(以AlpacaEval衡量)。本研究揭示了上下文学习与指令遵循能力间的潜在权衡,值得实际应用中关注。
原文摘要 · Abstract (English)
Instruct models, obtained from various instruction tuning or post-training steps, are commonly deemed superior and more usable than their base counterpart. While the model gains instruction following ability, instruction tuning may lead to forgetting the knowledge from pre-training or it may encourage the model to become overly conversational or verbose. This, in turn, can lead to degradation of in-context few-shot learning performance. In this work, we study the performance trajectory between base and instruct models by scaling down the strength of instruction-tuning via the partial adaption method. We show that, across several model families and model sizes, reducing the strength of instruction-tuning results in material improvement on a few-shot in-context learning benchmark covering a variety of classic natural language tasks. This comes at the cost of losing some degree of instruction following ability as measured by AlpacaEval. Our study shines light on the potential trade-off between in-context learning and instruction following abilities that is worth considering in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。