从数据特性出发,系统梳理联邦图学习的研究脉络与挑战。
A Comprehensive Data-centric Overview of Federated Graph Learning
- 按数据结构与分布特性分类,构建双层数据中心分类体系。
- 揭示数据利用策略如何应对分布式图数据的挑战。
- 适合关注图学习与联邦学习融合的科研人员参考。
在大数据应用时代,联邦图学习(FGL)已成为平衡去中心化数据持有者间集体智能优化与敏感信息保护的关键方案。现有综述多聚焦于联邦学习(FL)与图机器学习(GML)的结合,形成早期方法论分类,侧重模拟场景。然而,以数据为中心的视角——即从数据属性与使用方式系统审视FGL方法——尚未被充分整合,而这对于评估FGL研究如何应对数据约束以提升模型性能至关重要。本文提出两级数据中心分类:数据特征(依据数据集的结构与分布属性分类)与数据利用(分析训练流程与技术以应对核心数据挑战)。每级分类由三个正交标准构成,代表不同的数据配置。除分类体系外,本文还探讨了FGL与预训练大模型的融合、展示真实应用场景,并指出与图机器学习新兴趋势对齐的未来方向。
原文摘要 · Abstract (English)
In the era of big data applications, Federated Graph Learning (FGL) has emerged as a prominent solution that reconcile the tradeoff between optimizing the collective intelligence between decentralized datasets holders and preserving sensitive information to maximum. Existing FGL surveys have contributed meaningfully but largely focus on integrating Federated Learning (FL) and Graph Machine Learning (GML), resulting in early stage taxonomies that emphasis on methodology and simulated scenarios. Notably, a data centric perspective, which systematically examines FGL methods through the lens of data properties and usage, remains unadapted to reorganize FGL research, yet it is critical to assess how FGL studies manage to tackle data centric constraints to enhance model performances. This survey propose a two-level data centric taxonomy: Data Characteristics, which categorizes studies based on the structural and distributional properties of datasets used in FGL, and Data Utilization, which analyzes the training procedures and techniques employed to overcome key data centric challenges. Each taxonomy level is defined by three orthogonal criteria, each representing a distinct data centric configuration. Beyond taxonomy, this survey examines FGL integration with Pretrained Large Models, showcases realistic applications, and highlights future direction aligned with emerging trends in GML.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。