用大模型指导小模型对齐数据构建,提升微调效果
A Post-Training Enhanced Optimization Approach for Small Language Models
- 借大模型生成高质量对齐数据,增强多样性与准确性
- 在Qwen2-0.5B-Instruct上验证,性能显著优于传统微调方法
- 适合资源有限但需高性能小模型的场景
本文研究小语言模型的持续后训练优化方法,提出一种基于大模型数据引导的小模型对齐数据构建方法,以优化对齐数据的多样性和准确性。为验证该方法的有效性,以Qwen2-0.5B-Instruct模型作为小模型基线,使用所提方法构建的对齐数据集,开展多组实验对比,包括SFT(监督微调)、KTO(Kahneman-Tversky优化)、SFT-KTO两阶段后训练及模型权重融合实验。最终评估与分析表明,本文提出的持续后训练优化方法能显著提升小语言模型的性能。
原文摘要 · Abstract (English)
This paper delves into the continuous post-training optimization methods for small language models, and proposes a continuous post-training alignment data construction method for small language models. The core of this method is based on the data guidance of large models, optimizing the diversity and accuracy of alignment data. In addition, to verify the effectiveness of the methods in this paper, we used Qwen2-0.5B-Instruct model as the baseline model for small language models, using the alignment dataset constructed by our proposed method, we trained and compared several groups of experiments, including SFT (Supervised Fine Tuning) post-training experiment and KTO (Kahneman Tversky optimization) post-training experiment, as well as SFT-KTO two-stage post-training experiment and model weight fusion experiment. Finally, we evaluated and analyzed the performance of post-training models, and confirmed that the continuous post-training optimization method proposed by us can significantly improve the performance of small language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。