arXiv:2409.14655cs.DCcs.CR2024-09被引 2

通过自适应采样提升联邦图学习效率,大幅降低通信与计算开销。

Federated Graph Learning with Adaptive Importance-based Sampling

  • 基于图结构和优化动态的自适应重要节点采样
  • 测试准确率最高提升3.23%,通信与计算成本降低超90%
  • 适合大规模分布式图数据隐私保护场景

针对涉及分布式图数据集的隐私保护图学习任务,需采用基于联邦学习的GCN(FedGCN)训练。然而,现有方法在处理大规模图时面临计算与通信开销过高的问题,尤其因邻居数量爆炸式增长导致资源消耗巨大。现有图采样增强的FedGCN方法忽视图结构信息或优化动态,造成高方差与嵌入不准确。为此,本文提出联邦自适应重要性采样(FedAIS)方法,通过聚焦关键节点节省计算成本,并借助自适应历史嵌入同步减少通信开销。该方法联合考虑图结构异质性与优化动态,实现效率与精度的最优权衡。在五个真实世界图数据集上对五种先进基线进行广泛评估,结果显示FedAIS在保持相当或更高测试准确率(最高达3.23%提升)的同时,通信与计算成本分别降低91.77%和85.59%。

原文摘要 · Abstract (English)

For privacy-preserving graph learning tasks involving distributed graph datasets, federated learning (FL)-based GCN (FedGCN) training is required. A key challenge for FedGCN is scaling to large-scale graphs, which typically incurs high computation and communication costs when dealing with the explosively increasing number of neighbors. Existing graph sampling-enhanced FedGCN training approaches ignore graph structural information or dynamics of optimization, resulting in high variance and inaccurate node embeddings. To address this limitation, we propose the Federated Adaptive Importance-based Sampling (FedAIS) approach. It achieves substantial computational cost saving by focusing the limited resources on training important nodes, while reducing communication overhead via adaptive historical embedding synchronization. The proposed adaptive importance-based sampling method jointly considers the graph structural heterogeneity and the optimization dynamics to achieve optimal trade-off between efficiency and accuracy. Extensive evaluations against five state-of-the-art baselines on five real-world graph datasets show that FedAIS achieves comparable or up to 3.23% higher test accuracy, while saving communication and computation costs by 91.77% and 85.59%.

联邦学习图神经网络采样优化隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。