arXiv:2505.06911cs.LGcs.AI2025-05被引 2

解决多模态联邦学习中数据缺失问题,提升模型性能与隐私保护。

MMiC: Mitigating Modality Incompleteness in Clustered Federated Learning

  • 通过替换客户端模型部分参数缓解模态缺失影响。
  • 在多个数据集上优于现有方法,全局与个性化性能均提升。
  • 适合存在数据不完整或隐私限制的分布式多模态学习场景。

在大数据时代,数据挖掘对从海量复杂数据中发现隐藏模式至关重要。融合多模态数据源进一步增强了其潜力。多模态联邦学习(MFL)是一种分布式方法,可提升多模态学习的效率与质量,同时保障协作与隐私安全。然而,由于客户端间数据质量差异或隐私政策限制,模态缺失成为关键挑战。本文提出MMiC框架,用于在集群内缓解多模态联邦学习中的模态不完整问题。该框架通过替换集群内客户端模型的部分参数,减轻模态缺失的影响;利用巴赞夫权力指数优化客户端选择;并创新性地采用马科维茨投资组合理论动态控制全局聚合。大量实验表明,MMiC在包含模态缺失的多模态数据集上,持续优于现有联邦学习架构,在全局与个性化性能方面均表现更优,验证了所提方案的有效性。代码已开源:https://github.com/gotobcn8/MMiC。

原文摘要 · Abstract (English)

In the era of big data, data mining has become indispensable for uncovering hidden patterns and insights from vast and complex datasets. The integration of multimodal data sources further enhances its potential. Multimodal Federated Learning (MFL) is a distributed approach that enhances the efficiency and quality of multimodal learning, ensuring collaborative work and privacy protection. However, missing modalities pose a significant challenge in MFL, often due to data quality issues or privacy policies across the clients. In this work, we present MMiC, a framework for Mitigating Modality incompleteness in MFL within the Clusters. MMiC replaces partial parameters within client models inside clusters to mitigate the impact of missing modalities. Furthermore, it leverages the Banzhaf Power Index to optimize client selection under these conditions. Finally, MMiC employs an innovative approach to dynamically control global aggregation by utilizing Markovitz Portfolio Optimization. Extensive experiments demonstrate that MMiC consistently outperforms existing federated learning architectures in both global and personalized performance on multimodal datasets with missing modalities, confirming the effectiveness of our proposed solution. Our code is available at https://github.com/gotobcn8/MMiC.

联邦学习多模态数据缺失隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。