arXiv:2504.12025cs.LG2025-04被引 2

解决多模态联邦学习中数据异构与标签稀缺问题

FedEPA: Enhancing Personalization and Modality Alignment in Multimodal Federated Learning

  • 基于标注数据个性化聚合权重,缓解客户端数据差异
  • 无监督对齐跨模态特征,提升融合效果,准确率显著提升
  • 适合医疗、金融等多模态数据且标注少的隐私敏感场景

联邦学习(FL)可在保护隐私的前提下实现多方分布式模型训练,但现有系统大多假设客户端仅持有单模态数据,限制了实际应用。机构通常拥有多种模态数据,且标注数据有限,进一步制约性能。本文提出 FedEPA,一种新型多模态联邦学习框架。该方法采用个性化本地模型聚合策略,利用客户端的标注数据学习个性化聚合权重,缓解数据异构影响。同时提出无监督模态对齐策略:将多模态特征分解为对齐特征与上下文特征,通过对比学习对齐跨模态对齐特征,确保每模态内对齐特征与上下文特征独立,并促进上下文特征多样性。引入多模态特征融合策略生成联合嵌入。实验表明,在标注数据有限条件下,FedEPA 在多模态分类任务上显著优于现有联邦学习方法。

原文摘要 · Abstract (English)

Federated Learning (FL) enables decentralized model training across multiple parties while preserving privacy. However, most FL systems assume clients hold only unimodal data, limiting their real-world applicability, as institutions often possess multimodal data. Moreover, the lack of labeled data further constrains the performance of most FL methods. In this work, we propose FedEPA, a novel FL framework for multimodal learning. FedEPA employs a personalized local model aggregation strategy that leverages labeled data on clients to learn personalized aggregation weights, thereby alleviating the impact of data heterogeneity. We also propose an unsupervised modality alignment strategy that works effectively with limited labeled data. Specifically, we decompose multimodal features into aligned features and context features. We then employ contrastive learning to align the aligned features across modalities, ensure the independence between aligned features and context features within each modality, and promote the diversity of context features. A multimodal feature fusion strategy is introduced to obtain a joint embedding. The experimental results show that FedEPA significantly outperforms existing FL methods in multimodal classification tasks under limited labeled data conditions.

联邦学习多模态个性化无监督对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。