arXiv:2410.04287cs.LGcs.CY2024-10被引 8

揭示局部同质性如何导致图神经网络不公平,提出新基准与生成器。

Unveiling the Impact of Local Homophily on GNN Fairness: In-Depth Analysis and New Benchmarks

  • 将局部同质性偏差视为分布外问题,分析其对公平性的干扰机制。
  • 发现两种加剧不公平的因素:分布外距离与异质节点在同质图中存在。
  • 实测真实与合成数据集上公平性下降最高达30%,适合关注模型公平性研究者。

图神经网络(GNN)在同时存在同质性(同类连接)和异质性(异类连接)的图中常出现泛化能力下降,尤其当节点的局部同质性水平与全局水平差异显著时表现更差。这一问题在用户导向的应用中带来风险,因少数群体的同质性水平可能被忽视。尽管图神经网络的公平性已受关注,但局部同质性与公平性的关联仍不明确。本文突破全局同质性视角,深入探讨局部同质性如何引发不公平预测。首先将少数群体同质性偏差导致的公平性问题形式化为分布外(OOD)问题;随后进行理论分析,揭示局部同质性水平可改变不同敏感属性下的预测结果。此外,引入三个新的GNN公平性基准和一种新型半合成图生成器,用于实证研究此OOD问题。在广泛实验中发现,两个因素会加剧不公平:(a) 分布外距离,以及 (b) 异质节点存在于同质图中。当两者共现时,真实数据集公平性下降最高达24%,半合成数据集下降达30%。本工作通过理论洞察、实证分析与算法贡献,揭示了根植于图结构同质性信息的全新不公平来源。

原文摘要 · Abstract (English)

Graph Neural Networks (GNNs) often struggle to generalize when graphs exhibit both homophily (same-class connections) and heterophily (different-class connections). Specifically, GNNs tend to underperform for nodes with local homophily levels that differ significantly from the global homophily level. This issue poses a risk in user-centric applications where underrepresented homophily levels are present. Concurrently, fairness within GNNs has received substantial attention due to the potential amplification of biases via message passing. However, the connection between local homophily and fairness in GNNs remains underexplored. In this work, we move beyond global homophily and explore how local homophily levels can lead to unfair predictions. We begin by formalizing the challenge of fair predictions for underrepresented homophily levels as an out-of-distribution (OOD) problem. We then conduct a theoretical analysis that demonstrates how local homophily levels can alter predictions for differing sensitive attributes. We additionally introduce three new GNN fairness benchmarks, as well as a novel semi-synthetic graph generator, to empirically study the OOD problem. Across extensive analysis we find that two factors can promote unfairness: (a) OOD distance, and (b) heterophilous nodes situated in homophilous graphs. In cases where these two conditions are met, fairness drops by up to 24% on real world datasets, and 30% in semi-synthetic datasets. Together, our theoretical insights, empirical analysis, and algorithmic contributions unveil a previously overlooked source of unfairness rooted in the graph's homophily information.

图神经网络公平性同质性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。