梳理多模态联邦学习在三种主流范式下的挑战与方法
Multimodal Federated Learning: A Survey through the Lens of Different FL Paradigms
- 按水平、垂直、混合联邦范式分类分析多模态学习
- 揭示模态异质性、隐私异质性等独特挑战
- 适合关注隐私保护与多源数据融合的研究者
多模态联邦学习(MFL)结合了多模态信息互补与分布式训练的优势,提升下游推理性能并保护隐私。尽管关注度上升,当前尚无系统性分类框架来组织MFL在不同联邦学习(FL)范式下的研究。本文从三大主流FL范式——水平联邦学习(HFL)、垂直联邦学习(VFL)和混合联邦学习(Hybrid FL)出发,系统分析其问题建模、代表性算法及多模态数据带来的核心挑战。重点指出模态异质性、隐私异质性与通信效率低下等问题,在不同分布场景下表现各异,显著区别于传统单模态或非联邦学习场景。文章还讨论开放问题并提出未来研究方向。通过建立该分类体系,旨在揭示多模态数据在不同联邦设置下的新型挑战,为理解与推进MFL发展提供新视角。
原文摘要 · Abstract (English)
Multimodal Federated Learning (MFL) lies at the intersection of two pivotal research areas: leveraging complementary information from multiple modalities to improve downstream inference performance and enabling distributed training to enhance efficiency and preserve privacy. Despite the growing interest in MFL, there is currently no comprehensive taxonomy that organizes MFL through the lens of different Federated Learning (FL) paradigms. This perspective is important because multimodal data introduces distinct challenges across various FL settings. These challenges, including modality heterogeneity, privacy heterogeneity, and communication inefficiency, are fundamentally different from those encountered in traditional unimodal or non-FL scenarios. In this paper, we systematically examine MFL within the context of three major FL paradigms: horizontal FL (HFL), vertical FL (VFL), and hybrid FL. For each paradigm, we present the problem formulation, review representative training algorithms, and highlight the most prominent challenge introduced by multimodal data in distributed settings. We also discuss open challenges and provide insights for future research. By establishing this taxonomy, we aim to uncover the novel challenges posed by multimodal data from the perspective of different FL paradigms and to offer a new lens through which to understand and advance the development of MFL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。