大模型不一定是好老师,教师与学生模型的适配性更重要
Stronger Models are NOT Stronger Teachers for Instruction Tuning
- 提出兼容性调整奖励(CAR)衡量教师模型效果
- 实验证明大模型未必是小模型的好教师
- 适合优化指令微调中教师选择问题的研究者
指令微调已被广泛用于提升大语言模型(LLMs)遵循用户指令的能力,其效果高度依赖于微调所用的指令数据集。近年来,合成指令数据集因其经济高效成为提供多样化高质量指令的重要方案。然而,现有方法普遍假设更大的或更强的模型能作为更优的教师,因此直接使用这些模型生成合成指令的回答。本文挑战这一常见假设:通过在五种基础模型和二十个响应生成器上的广泛实验,我们发现更大更强的模型并不必然成为更优的教师。我们称此现象为「大模型悖论」。现有评估指标无法准确预测教师的有效性,因它们忽略了教师与待微调基模型之间的兼容性。为此,我们提出一种新指标——兼容性调整奖励(CAR),用于衡量响应生成器的效果。跨五种基础模型的实验表明,CAR显著优于几乎所有基线。
原文摘要 · Abstract (English)
Instruction tuning has been widely adopted to ensure large language models (LLMs) follow user instructions effectively. The resulting instruction-following capabilities of LLMs heavily rely on the instruction datasets used for tuning. Recently, synthetic instruction datasets have emerged as an economically viable solution to provide LLMs diverse and high-quality instructions. However, existing approaches typically assume that larger or stronger models are stronger teachers for instruction tuning, and hence simply adopt these models as response generators to the synthetic instructions. In this paper, we challenge this commonly-adopted assumption. Our extensive experiments across five base models and twenty response generators reveal that larger and stronger models are not necessarily stronger teachers of smaller models. We refer to this phenomenon as the Larger Models' Paradox. We observe that existing metrics cannot precisely predict the effectiveness of response generators since they ignore the compatibility between teachers and base models being fine-tuned. We thus develop a novel metric, named as Compatibility-Adjusted Reward (CAR) to measure the effectiveness of response generators. Our experiments across five base models demonstrate that CAR outperforms almost all baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。