arXiv:2607.23245cs.MMcs.LG2026-07中稿 · ACM Multimedia

通过结构迁移解决多模态联邦学习中缺失模态问题,提升协作效率。

FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities

论文配图:FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities
图 1 · 摘自论文原文
  • 利用冻结基础模型构建全局语义结构蓝图
  • 在非独立同分布数据下性能优于现有方法,通信开销降低
  • 适合缺模态场景下的高效多模态联邦学习

多模态联邦学习常面临任意模态缺失与非独立同分布数据分布问题,导致表示漂移并阻碍客户端间有效协作。现有方法依赖生成补全、外部辅助数据或孤立单模态训练,往往带来高昂通信与计算成本及潜在隐私风险。为此,我们提出FedTaste,一种面向缺失模态的拓扑感知结构迁移框架。不依赖脆弱的一阶特征对齐,而是聚焦更稳定的群体级语义关系。具体地,FedTaste利用冻结基础模型从完整模态客户端提取联合多模态拓扑,并由服务器整合为全局结构蓝图。针对缺失模态客户端,引入模态自适应结构提示与谱一致性正则化,实现轻量级分支特定适配,使局部部分表示与共享蓝图对齐。该方法避免显式模态补全,同时保持客户端间共享语义结构。大量实验表明,FedTaste在多个数据集和挑战性非独立同分布设置下持续取得优异性能,且相比现有方法显著降低通信开销。

原文摘要 · Abstract (English)

Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients. Existing methods typically rely on generative imputation, external auxiliary data, or isolated unimodal training to bridge modality gaps, often incurring substantial communication and computational costs as well as potential privacy risks. To address these limitations, we propose FedTaste, a parameter-efficient framework for topology-aware structural transfer in Multimodal Federated Learning with missing modalities. Instead of aligning fragile first-order features, FedTaste focuses on more stable group-level semantic relations. Specifically, FedTaste leverages frozen foundation models to extract a joint multimodal topology from full-modality clients, which is then consolidated by the server into a global structural blueprint. To adapt clients with missing modalities, we introduce Modality-Adaptive Structural Prompts together with spectral consistency regularization, enabling lightweight branch-specific adaptation that aligns local partial representations with the shared blueprint. In this way, FedTaste avoids explicit modality imputation while preserving shared semantic structure across clients. Extensive experiments demonstrate that FedTaste consistently achieves superior performance across multiple datasets and challenging Non-IID settings, while substantially reducing communication overhead compared with existing methods.

联邦学习多模态结构迁移缺模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。