arXiv:2505.01652cs.LGcs.AI2025-05中稿 · TMLR: https://open…

基于因果模型解决非独立同分布图数据的公平节点分类问题

Causally Fair Node Classification on Non-IID Graph Data

  • 提出MPVA框架,通过消息传递变分自编码器计算干预分布
  • 在半合成与真实数据上显著降低偏差,优于传统方法
  • 适用于社交网络等存在结构异质性的复杂图数据场景

公平机器学习旨在识别并缓解预测中针对种族、性别等人口属性的偏见。尽管已有研究将公平性扩展至图数据(如社交网络),但多数忽略了实例间的因果关系。本文从因果视角出发,针对经典公平学习常假设独立同分布(IID)数据的问题,研究节点因邻域结构差异而遵循不同因果机制的情形,此类情况违反经典结构性因果模型所需的不变性假设。基于网络结构性因果模型(NSCM)框架,提出消息传递变分自编码器(MPVA),用于计算因果公平节点分类的干预分布。在可分解性和图独立性两个条件下,建立了理论基础,形式化了如何通过构建结构表示恢复不变性,并在非IID设置下使用do-演算计算干预分布。在半合成与真实世界数据集上的实证评估表明,MPVA能有效逼近干预分布并减轻偏见,性能优于常规方法。研究结果展示了因果驱动公平性在复杂机器学习应用中的潜力,并为放松算法公平性中经典假设提供了新方向。

原文摘要 · Abstract (English)

Fair machine learning seeks to identify and mitigate biases in predictions against unfavorable populations characterized by demographic attributes, such as race and gender. Recent research has extended fairness to graph data, such as social networks, but many studies neglect the causal relationships among data instances. This paper addresses a prevalent challenge in many fair machine learning research, which typically assumes independent and identically distributed (IID) data, from the causal perspective. Specifically, this work targets the circumstance where nodes with different neighborhood structures follow different causal mechanisms, violating the invariance assumptions required for classical structural causal models and do-calculus. We base our research on the Network Structural Causal Model (NSCM) framework and develop a Message Passing Variational Autoencoder for Causal Inference (MPVA) to compute interventional distributions for causally fair node classification. We establish theoretical soundness under two conditions: Decomposability and Graph Independence. These conditions formalize when causal mechanism heterogeneity can be overcome by constructing a structural representation that restores invariance and facilitates the computation of interventional distributions using do-calculus in non-IID settings. Empirical evaluations on semi-synthetic and real-world datasets demonstrate that MPVA outperforms conventional methods by effectively approximating interventional distributions and mitigating bias. Our findings demonstrate the potential of causality-based fairness in complex ML applications and motivate future work on relaxing the classic assumptions in algorithmic fairness.

图神经网络因果推理公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。