arXiv:2506.14349cs.CYcs.IR2025-06被引 3

针对有限候选池的排名公平性问题,提出基于超几何分布的新评估框架。

hyperFA*IR: A hypergeometric approach to fair rankings with finite candidate pool

  • 用超几何分布建模无放回抽样,更贴合真实有限群体场景
  • 在小群体或高top-k比例时,比传统二项模型更准确反映统计特性
  • 可直接用于检测不公平排名,且支持设计有补偿作用的公平政策

排名算法在搜索、招聘等场景中影响重大,确保公平性尤为关键,尤其对数据中代表性不足的群体。现有方法多依赖比例代表制,但难以处理排名过程的随机性及候选池有限的问题。为此,本文提出 hyperFA*IR 框架,基于超几何分布建模从固定群体规模中无放回抽样,更真实地刻画实际场景。该方法在 top-$k$ 选择占比大或保护群体较小时,能更准确评估公平性。通过理论与实证对比,证明其优于常用的独立抽样二项模型。进一步提出基于蒙特卡洛的高效算法,无需复杂调参即可检测不公平排名,并可引入权重实现有补偿性的公平政策。

原文摘要 · Abstract (English)

Ranking algorithms play a pivotal role in decision-making processes across diverse domains, from search engines to job applications. When rankings directly impact individuals, ensuring fairness becomes essential, particularly for groups that are marginalised or misrepresented in the data. Most of the existing group fairness frameworks often rely on ensuring proportional representation of protected groups. However, these approaches face limitations in accounting for the stochastic nature of ranking processes or the finite size of candidate pools. To this end, we present hyperFA*IR, a framework for assessing and enforcing fairness in rankings drawn from a finite set of candidates. It relies on a generative process based on the hypergeometric distribution, which models real-world scenarios by sampling without replacement from fixed group sizes. This approach improves fairness assessment when top-$k$ selections are large relative to the pool or when protected groups are small. We compare our approach to the widely used binomial model, which treats each draw as independent with fixed probability, and demonstrate$-$both analytically and empirically$-$that our method more accurately reproduces the statistical properties of sampling from a finite population. To operationalise this framework, we propose a Monte Carlo-based algorithm that efficiently detects unfair rankings by avoiding computationally expensive parameter tuning. Finally, we adapt our generative approach to define affirmative action policies by introducing weights into the sampling process.

排名公平超几何分布有限池公平算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。