解决跨客户端多模态图学习中的异质性难题,提升协作性能。
Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity

- 分两阶段训练:先全局预训练,再独立微调应对任务差异。
- 引入拓扑感知的跨模态路由机制,提升异构数据对齐效果。
- 适合多模态图数据分散、模态不一致的工业场景应用。
多模态属性图(MAGs)在多个领域广泛应用,其节点包含跨模态的异质语义内容,边则编码关系依赖。联邦多模态图学习(FMGL)将联邦图学习扩展至MAGs,实现去中心化MAGs间的协作优化而无需暴露原始数据。然而,直接套用现有联邦图学习方法在FMGL中表现不佳,因其无法应对多维度异质性:包括客户端目标差异带来的任务异质性、模态质量与语义域差异导致的模态异质性,以及拓扑模式不同且跨模态相关性低引发的拓扑异质性。为此,本文提出首个系统性的FMGL算法——拓扑感知跨模态路由联邦学习(FedTCR)。针对任务异质性,采用两阶段范式:联邦任务无关预训练后进行独立任务微调;为联合处理模态与拓扑异质性,提出拓扑感知跨模态路由机制:各客户端基于图结构,通过拓扑感知重要性加权聚合,提炼出模态特定的紧凑原型;服务器评估跨客户端跨模态原型间关系,路由有信息量的作为对比参考,驱动三层次跨模态对比学习,同步实现跨客户端模态对齐并保持区分能力。在7个领域的实验表明,FedTCR在以图为中心和以模态为中心的任务上均优于现有最先进基线。
原文摘要 · Abstract (English)
Multimodal-attributed graphs (MAGs), where nodes carry heterogeneous semantic content across multiple modalities while edges encode relational dependencies, have been widely adopted across diverse domains. Federated multimodal graph learning (FMGL) extends federated graph learning (FGL) to MAGs, enabling collaborative optimization across decentralized MAGs without exposing raw data. However, naively applying existing FGL methods to FMGL is insufficient, as they fail to navigate the multifaceted heterogeneity inherent in decentralized MAGs, including task heterogeneity across diverse client objectives, modality heterogeneity from discrepant modality quality and semantic domains, and topology heterogeneity arising from divergent topological patterns with low cross-modality correlation. To address these challenges, we propose Federated multimodal graph learning with Topology-aware Cross-modal Routing (FedTCR), the first systematic algorithm designed for FMGL. To handle task heterogeneity, FedTCR employs a two-stage paradigm that comprises federated task-agnostic pre-training followed by isolated task-oriented fine-tuning. To jointly address modality and topology heterogeneity, FedTCR introduces a topology-aware cross-modal routing mechanism. Concretely, each client distills modality-specific knowledge into compact prototypes via topology-aware importance-weighted aggregation informed by graph structure; the server then evaluates cross-client cross-modal relationships among these structure-informed prototypes and routes informative ones as contrastive references, driving a tri-level cross-modal contrastive learning scheme that jointly aligns cross-client modalities while preserving discrimination. Experiments across 7 domains demonstrate that FedTCR outperforms state-of-the-art baselines on both graph-centric and modality-centric tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。