通过隐空间共识实现高效多模态联邦学习,提升跨模态一致性与抗噪能力。
Communication-Efficient and Robust Multi-Modal Federated Learning via Latent-Space Consensus
- 用可学习投影矩阵压缩多模态数据为隐空间表示
- 隐空间正则化使不同客户端表示对齐,准确率接近最优
- 适合数据异构、通信受限的多模态联邦场景
联邦学习(FL)可在不共享原始数据的前提下实现分布式设备协同训练,但将其应用于多模态场景面临显著挑战。客户端通常拥有异构的模态和模型结构,难以高效对齐特征空间,同时兼顾隐私保护与通信开销。为此,我们提出CoMFed——一种通信高效的多模态联邦学习框架,采用可学习投影矩阵生成压缩的隐空间表示,并通过隐空间正则化对齐各客户端的表示,提升跨模态一致性与对异常值的鲁棒性。在人体活动识别基准测试中,CoMFed实现了具有竞争力的准确率,且通信与计算开销极低。
原文摘要 · Abstract (English)
Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, but applying FL to multi-modal settings introduces significant challenges. Clients typically possess heterogeneous modalities and model architectures, making it difficult to align feature spaces efficiently while preserving privacy and minimizing communication costs. To address this, we introduce CoMFed, a Communication-Efficient Multi-Modal Federated Learning framework that uses learnable projection matrices to generate compressed latent representations. A latent-space regularizer aligns these representations across clients, improving cross-modal consistency and robustness to outliers. Experiments on human activity recognition benchmarks show that CoMFed achieves competitive accuracy with minimal overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。