arXiv:2411.18607cs.LG2024-11被引 13

用联邦学习视角解释模型合并技术,提升多任务模型融合效果

Task Arithmetic Through The Lens Of One-Shot Federated Learning

  • 将模型合并视为单次联邦学习问题,揭示其数学本质
  • 发现数据和训练异质性是影响融合效果的关键因素
  • 借鉴联邦学习算法,显著提升模型合并性能,适合模型集成研究者

Task Arithmetic 是一种在权重空间通过简单算术操作合并多个模型能力的技术,无需额外微调或原始训练数据。然而,其成功的影响因素尚不明确。本文将多任务学习中的 Task Arithmetic 视为单次联邦学习问题,证明其数学上等价于联邦学习中常用的 Federated Averaging (FedAvg) 算法。基于 FedAvg 的成熟理论,我们识别出影响任务合并性能的两个关键因素:数据异质性和训练异质性。为缓解这些问题,我们引入并适配了联邦学习中的多种算法以增强 Task Arithmetic 效果。实验表明,应用这些改进方法可显著提升合并模型的性能。本工作连接了 Task Arithmetic 与联邦学习,提供了新的理论视角与更优的模型合并实践方法。

原文摘要 · Abstract (English)

Task Arithmetic is a model merging technique that enables the combination of multiple models' capabilities into a single model through simple arithmetic in the weight space, without the need for additional fine-tuning or access to the original training data. However, the factors that determine the success of Task Arithmetic remain unclear. In this paper, we examine Task Arithmetic for multi-task learning by framing it as a one-shot Federated Learning problem. We demonstrate that Task Arithmetic is mathematically equivalent to the commonly used algorithm in Federated Learning, called Federated Averaging (FedAvg). By leveraging well-established theoretical results from FedAvg, we identify two key factors that impact the performance of Task Arithmetic: data heterogeneity and training heterogeneity. To mitigate these challenges, we adapt several algorithms from Federated Learning to improve the effectiveness of Task Arithmetic. Our experiments demonstrate that applying these algorithms can often significantly boost performance of the merged model compared to the original Task Arithmetic approach. This work bridges Task Arithmetic and Federated Learning, offering new theoretical perspectives on Task Arithmetic and improved practical methodologies for model merging.

模型融合联邦学习多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。