arXiv:2509.12396cs.SIcs.LG2025-09

揭示网络嵌入中因图结构导致的信息丢失与群体差异

Information Loss and Disparate Effects in Network Embeddings

  • 基于随机块模型分析嵌入的渐近行为,发现信息损失与图密度和同质性相关
  • 不同图结构可能生成相同嵌入,小而稀疏社区受影响更严重
  • 解释了下游任务如链接预测在少数群体中误差更高的现象

大量研究关注网络嵌入中的公平性干预,但对其基线行为了解有限。本文探究:无公平性干预的基线嵌入在表示层面如何产生群体差异?通过分析随机块模型(SBM)图上的低维嵌入渐近行为,我们刻画了信息丢失的精确条件,表明信息损失程度直接依赖于图的密度和同质性。值得注意的是,不同图结构在极限下可能产生完全相同的嵌入,且这种不可逆性对小而稀疏的社区影响尤为严重。结果导致简单下游任务(如链接预测)在这些社区中出现更高错误率,有助于解释实践中广泛观察到的不公平现象。

原文摘要 · Abstract (English)

An extensive line of work studies fairness interventions for network embeddings, but less is known about their baseline behavior. In this work, we ask: how do baseline embeddings (without fairness interventions) produce disparate effects at the representation level? We analyze the asymptotic behavior of low-dimensional embeddings on stochastic block model (SBM) graphs, which encode both homophily and group structure. We characterize exact conditions under which embeddings cause information loss, showing that the amount of information loss depends directly on the graph's density and assortativity. Notably, very different graphs can produce identical embeddings in the limit, and this non-invertibility disproportionately affects smaller and sparser communities. As a result, simple downstream tasks, such as link prediction, introduce higher error rates for these communities, helping explain disparities widely observed in practice.

网络嵌入公平性信息损失图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。