解决多模态联邦学习中数据缺失不一致的问题,让不同客户端的提示词能协同优化。
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
- 设计客户端与服务器协同优化的提示词调优机制,应对多模态数据缺失差异。
- 在多个多模态基准数据集上超越现有最优方法,提升模型泛化能力。
- 适合研究联邦学习与多模态大模型融合的开发者和研究人员。
本文提出一种通用的联邦提示调优框架,适用于本地数据为多模态且输入层存在不同缺失模式的实际场景。该框架弥合了传统联邦学习与集中式多模态提示调优之间的差距。关键挑战在于不同客户端间编码相似缺失分布的提示指令缺乏语义对齐。为此,我们设计了专门的客户端调优与服务器聚合方案,同步优化、对齐并聚合跨客户端和多模态的提示调优指令,使提示指令能够互补并有效结合。在多个多模态基准数据集上的广泛评估表明,本方法持续优于当前最优(SOTA)基线。
原文摘要 · Abstract (English)
This paper introduces a generalized federated prompt-tuning framework for practical scenarios where local datasets are multi-modal and exhibit different distributional patterns of missing features at the input level. The proposed framework bridges the gap between federated learning and multi-modal prompt-tuning which have traditionally focused on either uni-modal or centralized data. A key challenge in this setting arises from the lack of semantic alignment between prompt instructions that encode similar distributional patterns of missing data across different clients. To address this, our framework introduces specialized client-tuning and server-aggregation designs that simultaneously optimize, align, and aggregate prompt-tuning instructions across clients and data modalities. This allows prompt instructions to complement one another and be combined effectively. Extensive evaluations on diverse multimodal benchmark datasets demonstrate that our work consistently outperforms state-of-the-art (SOTA) baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。