提出F³OCUS方法,优化联邦学习中视觉语言模型的高效微调。
F$^3$OCUS -- Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristics
- 基于神经正切核主特征值设计客户端层重要性评分
- 联合优化层重要性与跨客户端多样性,提升微调效果
- 在医疗多模态数据上验证,支持多种模型架构
在资源受限的客户端设备上进行联邦学习时,高效训练大型视觉-语言模型(VLMs)需采用参数高效微调(PEFT)策略。本文探讨了两个关键因素:客户端特定的层重要性得分(用于选择最需微调的模型层)和客户端间层多样性得分(促进各客户端间不同层的选择)。我们首先理论推导并利用层间神经正切核的主特征值大小作为客户端层重要性评分的有效指标。随后提出新型层更新策略F³OCUS,通过服务器端无数据、多目标、元启发式优化,联合优化重要性与多样性。我们评估了5种元启发式算法在选择模型层和适配器层方面的有效性。此外,我们发布了新的MedVQA-FL数据集,包含707,962个VQA三元组及9个模态特定客户端,并在6个视觉-语言联邦学习任务设置下,使用58个医学图像数据集和4种不同规模的VLM架构,进行了超过10,000次客户端级实验,充分验证了该方法的有效性。
原文摘要 · Abstract (English)
Effective training of large Vision-Language Models (VLMs) on resource-constrained client devices in Federated Learning (FL) requires the usage of parameter-efficient fine-tuning (PEFT) strategies. To this end, we demonstrate the impact of two factors \textit{viz.}, client-specific layer importance score that selects the most important VLM layers for fine-tuning and inter-client layer diversity score that encourages diverse layer selection across clients for optimal VLM layer selection. We first theoretically motivate and leverage the principal eigenvalue magnitude of layerwise Neural Tangent Kernels and show its effectiveness as client-specific layer importance score. Next, we propose a novel layer updating strategy dubbed F$^3$OCUS that jointly optimizes the layer importance and diversity factors by employing a data-free, multi-objective, meta-heuristic optimization on the server. We explore 5 different meta-heuristic algorithms and compare their effectiveness for selecting model layers and adapter layers towards PEFT-FL. Furthermore, we release a new MedVQA-FL dataset involving overall 707,962 VQA triplets and 9 modality-specific clients and utilize it to train and evaluate our method. Overall, we conduct more than 10,000 client-level experiments on 6 Vision-Language FL task settings involving 58 medical image datasets and 4 different VLM architectures of varying sizes to demonstrate the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。