解决多模态联邦学习中的模态差异与个性化难题
FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation
- 客户端用对抗对齐和自适应融合提升跨模态表示
- 服务器通过相似性引导的集成蒸馏实现全局模型优化
- 在异构数据下兼顾性能与效率,适合隐私敏感场景
多模态联邦学习(MFL)使具有异构数据模态的客户端可在不共享原始数据的情况下协同训练模型,提供一种保护隐私并利用跨模态互补信息的框架。然而,现有方法常忽视个性化客户端表现,难以应对模态/任务差异及模型异构性。为此,我们提出统一的FedAFD框架,增强客户端与服务器的学习能力。客户端采用双层对抗对齐策略,对齐模态内与模态间的局部与全局表示,缓解模态与任务差距;设计粒度感知融合模块,自适应整合全局知识到个性化特征中。服务器侧针对模型异构性,提出基于相似性的集成蒸馏机制,基于共享公共数据上的特征相似性聚合客户端表示,并将融合知识蒸馏至全局模型。在IID与非IID设置下的大量实验表明,FedAFD在客户端与服务器端均实现更优性能与效率。
原文摘要 · Abstract (English)
Multimodal Federated Learning (MFL) enables clients with heterogeneous data modalities to collaboratively train models without sharing raw data, offering a privacy-preserving framework that leverages complementary cross-modal information. However, existing methods often overlook personalized client performance and struggle with modality/task discrepancies, as well as model heterogeneity. To address these challenges, we propose FedAFD, a unified MFL framework that enhances client and server learning. On the client side, we introduce a bi-level adversarial alignment strategy to align local and global representations within and across modalities, mitigating modality and task gaps. We further design a granularity-aware fusion module to integrate global knowledge into the personalized features adaptively. On the server side, to handle model heterogeneity, we propose a similarity-guided ensemble distillation mechanism that aggregates client representations on shared public data based on feature similarity and distills the fused knowledge into the global model. Extensive experiments conducted under both IID and non-IID settings demonstrate that FedAFD achieves superior performance and efficiency for both the client and the server.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。