无需数据的自适应模型融合,有效缓解任务干扰问题
FroM: Frobenius Norm-Based Data-Free Adaptive Model Merging
- 基于弗罗贝尼乌斯范数动态调整参数融合权重
- 在多种微调场景下均优于基线方法,无须训练数据
- 适合参数高效微调和多模型知识融合场景
随着大语言模型的发展,微调已成为通过注入领域知识提升特定场景性能的有效方法。在此背景下,模型融合技术通过合并多个微调模型的参数,实现知识融合。然而,传统方法在融合全量微调模型时常出现任务干扰,该问题在参数高效微调场景中尤为显著。本文改进了RegMean方法,间接利用训练数据近似线性层在融合前后的输出。提出一种名为FroM的自适应融合方法,直接使用弗罗贝尼乌斯范数测量模型参数,无需任何训练数据。通过引入额外超参数进行控制,FroM在多种微调场景下均优于基线方法,有效缓解了任务干扰问题。
原文摘要 · Abstract (English)
With the development of large language models, fine-tuning has emerged as an effective method to enhance performance in specific scenarios by injecting domain-specific knowledge. In this context, model merging techniques provide a solution for fusing knowledge from multiple fine-tuning models by combining their parameters. However, traditional methods often encounter task interference when merging full fine-tuning models, and this problem becomes even more evident in parameter-efficient fine-tuning scenarios. In this paper, we introduce an improvement to the RegMean method, which indirectly leverages the training data to approximate the outputs of the linear layers before and after merging. We propose an adaptive merging method called FroM, which directly measures the model parameters using the Frobenius norm, without any training data. By introducing an additional hyperparameter for control, FroM outperforms baseline methods across various fine-tuning scenarios, alleviating the task interference problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。