提出细粒度性别偏见评估与缓解方法,提升推荐系统公平性。
Unmasking Gender Bias in Recommendation Systems and Enhancing Category-Aware Fairness
- 按推荐类别(如电影类型)量化性别偏见,突破群体平均评估局限。
- 引入类别感知正则项后,各类别推荐公平性显著提升,性能损失小。
- 适用于关注推荐系统伦理、需优化性别公平性的研究与工程团队。
推荐系统已深度融入日常生活,用于发现新电影、社交交友及职位匹配等。由于其重要性,必须确保推荐结果不受社会刻板印象影响。现有公平性评估多基于敏感群体的性能对比,未能捕捉细微差异。本文提出一套全面的性别偏见量化指标,强调在更细粒度的项目类别层面(如电影类型)评估公平性的重要性。进一步地,将类别感知公平性指标作为正则项加入训练过程,可有效降低模型输出中的偏见。我们在三个真实数据集上,使用五种基线模型和两种主流公平性增强模型进行实验,验证了所提指标能提供比以往更深入的偏见洞察。结果表明,引入正则项后,各类别推荐公平性显著改善,整体推荐性能无明显下降。
原文摘要 · Abstract (English)
Recommendation systems are now an integral part of our daily lives. We rely on them for tasks such as discovering new movies, finding friends on social media, and connecting job seekers with relevant opportunities. Given their vital role, we must ensure these recommendations are free from societal stereotypes. Therefore, evaluating and addressing such biases in recommendation systems is crucial. Previous work evaluating the fairness of recommended items fails to capture certain nuances as they mainly focus on comparing performance metrics for different sensitive groups. In this paper, we introduce a set of comprehensive metrics for quantifying gender bias in recommendations. Specifically, we show the importance of evaluating fairness on a more granular level, which can be achieved using our metrics to capture gender bias using categories of recommended items like genres for movies. Furthermore, we show that employing a category-aware fairness metric as a regularization term along with the main recommendation loss during training can help effectively minimize bias in the models' output. We experiment on three real-world datasets, using five baseline models alongside two popular fairness-aware models, to show the effectiveness of our metrics in evaluating gender bias. Our metrics help provide an enhanced insight into bias in recommended items compared to previous metrics. Additionally, our results demonstrate how incorporating our regularization term significantly improves the fairness in recommendations for different categories without substantial degradation in overall recommendation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。