不加算力和数据,用外推合并让大模型持续变强
Extrapolation Merging: Keep Improving With Extrapolation and Merging
- 用外推法为模型合并提供明确优化方向
- 7个任务上均实现微调后性能提升
- 适合资源受限下持续优化大模型的场景
大型语言模型(LLMs)需通过指令微调完成下游任务,但该阶段仍需大量计算资源与标注数据,缺乏无需额外算力与数据即可提升性能的新范式。模型合并旨在通过融合不同模型参数来增强性能,但合并过程缺乏明确优化方向,难以保证效果提升。本文首次验证了指令微调阶段外推法的有效性,并提出一种名为外推合并(Extrapolation Merging)的新范式,可在不增加计算资源与数据的前提下持续提升模型性能。通过外推方法为合并过程提供清晰方向,实现局部优化搜索,从而增强合并后模型表现。我们在七个不同任务上进行了实验,结果表明,该方法在微调后能持续提升模型性能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) require instruction fine-tuning to perform different downstream tasks. However, the instruction fine-tuning phase still demands significant computational resources and labeled data, lacking a paradigm that can improve model performance without additional computational power and data. Model merging aims to enhance performance by combining the parameters of different models, but the lack of a clear optimization direction during the merging process does not always guarantee improved performance. In this paper, we attempt to provide a clear optimization direction for model merging. We first validate the effectiveness of the model extrapolation method during the instruction fine-tuning phase. Then, we propose Extrapolation Merging, a paradigm that can continue improving model performance without requiring extra computational resources or data. Using the extrapolation method, we provide a clear direction for model merging, achieving local optimization search, and consequently enhancing the merged model's performance. We conduct experiments on seven different tasks, and the results show that our method can consistently improve the model's performance after fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。