针对图联邦学习中的同质性差异问题,提出新方法提升协作效果。
Homophily Heterogeneity Matters in Graph Federated Learning: A Spectrum Sharing and Complementing Perspective
- 基于图谱特性设计联邦学习框架,实现低频信息共享与高频信息互补。
- 在异质图数据上平均性能超越第二名3.28%,显著提升模型表现。
- 适合处理同质性差异大的跨客户端图学习任务,尤其适用于异质图场景。
由于异质性是图联邦学习中的根本挑战,现有方法主要关注节点特征异质性和结构异质性,却忽略了关键的同质性异质性——即不同客户端图数据间的同质性水平存在显著差异。同质性水平指连接同类节点的边所占比例。因适应本地同质性,各客户端的局部模型捕捉到不一致的谱特性,严重削弱协作效率。具体而言,高同质性图上的模型倾向于捕获低频信息,而低同质性图上的模型则更关注高频信息。为此,本文引入谱图神经网络,提出新型联邦学习方法FedGSP,通过挖掘图谱特性来应对同质性异质性。一方面,该方法使客户端共享通用谱特性(即低频信息),实现协同增益;另一方面,基于理论发现,允许客户端互补缺乏的非通用谱特性(即高频信息),从而获得额外信息增益。在六个同质图和五个异质图数据集上,涵盖非重叠与重叠设置的大量实验验证了该方法优于十一种先进方法。特别地,在所有异质图数据集上,其平均性能超过第二优方法3.28%。
原文摘要 · Abstract (English)
Since heterogeneity presents a fundamental challenge in graph federated learning, many existing methods are proposed to deal with node feature heterogeneity and structure heterogeneity. However, they overlook the critical homophily heterogeneity, which refers to the substantial variation in homophily levels across graph data from different clients. The homophily level represents the proportion of edges connecting nodes that belong to the same class. Due to adapting to their local homophily, local models capture inconsistent spectral properties across different clients, significantly reducing the effectiveness of collaboration. Specifically, local models trained on graphs with high homophily tend to capture low-frequency information, whereas local models trained on graphs with low homophily tend to capture high-frequency information. To effectively deal with homophily heterophily, we introduce the spectral Graph Neural Network (GNN) and propose a novel Federated learning method by mining Graph Spectral Properties (FedGSP). On one hand, our proposed FedGSP enables clients to share generic spectral properties (i.e., low-frequency information), allowing all clients to benefit through collaboration. On the other hand, inspired by our theoretical findings, our proposed FedGSP allows clients to complement non-generic spectral properties by acquiring the spectral properties they lack (i.e., high-frequency information), thereby obtaining additional information gain. Extensive experiments conducted on six homophilic and five heterophilic graph datasets, across both non-overlapping and overlapping settings, validate the superiority of our method over eleven state-of-the-art methods. Notably, our FedGSP outperforms the second-best method by an average margin of 3.28% on all heterophilic datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。