现有打分方法无法兼顾公平与效率,需改进后处理机制
Scoring Is Not Enough: Addressing Gaps in Utility-fairness Trade-offs for Ranking

- 用反例证明打分函数在公平-效用权衡上存在根本缺陷
- 半贪婪后处理可显著提升公平与效用的综合表现
- 适合关注推荐系统公平性的研究人员和工程师
打分函数用于衡量单个文档的相关性,在现代信息检索或推荐系统中常通过数据学习得到,以最大化查询或用户需求下的系统效用。随着算法公平性受关注,研究者尝试设计同时权衡公平与效用的打分方法。本文通过一系列反例,证明无论采用确定性或随机性打分函数,也无论在单个查询或跨多个查询层面衡量公平性,打分方法均无法实现所有公平-效用权衡。正面结果表明,半贪婪后处理方法可在可接受计算成本下,接近穷举后处理的理想效果,显著改善整体权衡性能。
原文摘要 · Abstract (English)
Scoring functions are used to represent the relevance of individual documents. In modern information retrieval or recommendation systems, they are often learned from data and play a pivotal role in ranking sets of documents or items in a way that maximizes utility to a query or user. With the recent interest in algorithmic fairness, the success of scoring has naturally led to methods that learn scores that simultaneously trade off fairness and utility. In this work, we show that in stark contrast with utility-centric objectives, scoring is sub-optimal in achieving all utility-fairness trade-offs. We establish this with a series of counter-examples with a generic fairness formulation. We show that the issue persists whether we have a deterministic scoring function or a randomized one, or whether we measure fairness at the scope of a single query or across multiple queries. On the positive side, we empirically demonstrate that semi-greedy post-processing has the potential to achieve much better trade-offs, often approaching the ideal of exhaustive post-processing in a tractable way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。