提出视觉语言模型联邦学习的理论框架,优化个性化与泛化平衡。
Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and Method

- 用特征学习理论分析提示词在联邦学习中的表现
- 证明任务相关系数占比越高性能越好
- 设计全局+局部提示组合,提升效果并求出最优权重
将预训练的视觉-语言基础模型(如CLIP)引入联邦学习,可增强跨任务泛化能力。当前多采用提示学习降低通信与计算开销,即提示词联邦学习。然而,该方法缺乏理论分析。本文基于特征学习理论,构建提示词联邦学习的理论分析框架,监测信号学习与噪声记忆的演变过程,证明性能由任务相关系数与无关系数的比率决定。类比投资组合中收益与风险的关系,借鉴资产组合优化思想,引入全局提示与本地提示构成提示组合,在保持任务相关性的同时降低无关干扰。由此实现泛化与个性化的平衡,并推导出最优混合系数。理论结论得到实验验证。
原文摘要 · Abstract (English)
Integrating pretrained vision-language foundation models like CLIP into federated learning has attracted significant attention for enhancing generalization across diverse tasks. Typically, federated learning of vision-language models employs prompt learning to reduce communication and computational costs, i.e., prompt-based federated learning. However, there is limited theoretical analysis to understand the performance of prompt-based federated learning. In this work, we construct a theoretical analysis framework for prompt-based federated learning via feature learning theory. Specifically, we monitor the evolution of signal learning and noise memorization in prompt-based federated learning, demonstrating that performance can be assessed by the ratio of task-relevant to task-irrelevant coefficients. Furthermore, we draw an analogy between income and risk in portfolio optimization and the task-relevant and task-irrelevant terms in feature learning. Leveraging inspiration from portfolio optimization that combining two independent assets will maintain the income while reducing the risk, we introduce two prompts: global prompt and local prompt to construct a prompt portfolio to balance the generalization and personalization. Consequently, we showed the performance advantage of the prompt portfolio and derived the optimal mixing coefficient. These theoretical claims have been further supported by empirical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。