检验推荐模型公平性:代表层面的公平不等于推荐结果的公平
Exploring How Fair Model Representations Relate to Fair Recommendations
- 用排名推荐结果直接评估性别等信息泄露程度
- 实验证明优化表示公平性能提升推荐一致性
- 适合关注推荐系统公平性评估方法的研究者
近年来推荐系统研究中,一种常见公平性目标是减少模型表示中编码的人口统计信息。通常通过判断给定模型表示后能否准确分类人口属性来评估,隐含假设是该指标能反映推荐一致性(即不同用户获得相似推荐的程度)。本文挑战这一假设,比较了模型表示中的信息量与多种推荐差异度量之间的关系。提出两种新方法,基于排序后的推荐结果来衡量人口信息可分类性。在真实数据集及多个合成数据集上对多种模型进行广泛测试的结果表明:优化表示公平性确实有助于提升推荐一致性,但仅在表示层面评估公平性并不能有效反映模型间的推荐一致性差异。同时,通过在具有不同特性的生成数据集上评估多种模型,深入揭示了推荐层面公平性指标的表现规律。
原文摘要 · Abstract (English)
One of the many fairness definitions pursued in recent recommender system research targets mitigating demographic information encoded in model representations. Models optimized for this definition are typically evaluated on how well demographic attributes can be classified given model representations, with the (implicit) assumption that this measure accurately reflects \textit{recommendation parity}, i.e., how similar recommendations given to different users are. We challenge this assumption by comparing the amount of demographic information encoded in representations with various measures of how the recommendations differ. We propose two new approaches for measuring how well demographic information can be classified given ranked recommendations. Our results from extensive testing of multiple models on one real and multiple synthetically generated datasets indicate that optimizing for fair representations positively affects recommendation parity, but also that evaluation at the representation level is not a good proxy for measuring this effect when comparing models. We also provide extensive insight into how recommendation-level fairness metrics behave for various models by evaluating their performances on numerous generated datasets with different properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。