arXiv:2606.01856cs.DCcs.AI2026-06

解决多模态联邦学习中模态竞争问题,提升模型性能与通信效率

Boosting Multimodal Federated Learning via Chained Modality Optimization

论文配图:Boosting Multimodal Federated Learning via Chained Modality Optimization
图 1 · 摘自论文原文
  • 将多模态训练设计为分阶段优化,每模态独立优化避免压制
  • 在多个基准上实现更高预测性能,通信频率更低
  • 适合数据异构、模态缺失的隐私保护协同学习场景

多模态联邦学习(MMFL)支持在数据异构且模态不全的分布式客户端间进行隐私保护协作学习。然而,现有方法多将多模态训练视为联合优化问题,忽视了关键瓶颈:模态竞争,即主导模态抑制弱模态,导致全局模型次优。为此,我们提出 FedMChain,一种平衡的多模态联邦学习框架,将联邦多模态训练组织为一系列模态专属阶段。该分阶段设计为每个模态在多模态客户端上提供专用本地优化窗口,缓解模态竞争,并通过误差补偿正则项促进跨模态互补性。服务器端采用稀疏符号引导聚合策略,利用方向符号一致性实现稳健的模态内聚合,避免破坏性平均,支持更低频同步以降低通信开销。在多个多模态基准上的大量实验表明,FedMChain持续提升预测性能,同时通信频率低于基线方法。

原文摘要 · Abstract (English)

Multimodal Federated Learning (MMFL) enables privacy-preserving collaborative learning across decentralized clients with heterogeneous data and modality availability. However, most existing MMFL methods cast multimodal training as a joint optimization problem, overlooking a key bottleneck: modality competition, where dominant modalities suppress weaker ones and lead to suboptimal global models. To address this, we propose FedMChain, a balanced MMFL framework that structures federated multimodal training as a chain of modality-wise phases. This phase-wise design gives each modality a dedicated local optimization window on multimodal clients to mitigate modality competition, and further promotes cross-modal complementarity via an error-compensated regularizer. On the server side, we employ a sparse sign-guided aggregation strategy that leverages directional sign agreement for robust intra-modality aggregation, avoids destructive averaging, and supports less frequent synchronization to reduce communication overhead. Extensive experiments on multimodal benchmarks demonstrate that FedMChain consistently improves predictive performance while requiring less frequent communication than baselines.

联邦学习多模态通信效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。