arXiv:2501.04420cs.IR2025-01

攻击者可利用性别刻板印象从观影数据中推断用户性别,暴露隐私风险。

A Closer Look on Gender Stereotypes in Movie Recommender Systems and Their Implications with Privacy

  • 通过用户调研识别电影类型与性别的刻板关联
  • 四种算法结合刻板印象使性别推断准确率高于仅用反馈数据
  • 适用于关注推荐系统隐私与偏见的研究者

电影推荐系统通常基于用户反馈提供个性化推荐以提升收益。本研究通过特定攻击场景探究性别刻板印象对系统的负面影响:攻击者利用人们对电影偏好的性别刻板印象,结合用户公开或系统内观测的反馈数据,推断用户性别这一私密属性。研究分两阶段进行:第一阶段通过630名参与者用户研究,识别出与性别相关的电影类型刻板印象;第二阶段应用四种推理算法,结合第一阶段结果与用户反馈数据,验证其在性别推断中的有效性。实验结果显示,融合刻板印象的算法显著优于仅依赖反馈数据的方法。研究还使用MovieLens 1M和Yahoo!Movie两个主流数据集进行评估,量化了性别刻板印象在数字计算科学中的广泛影响。详细实验信息见GitHub:https://github.com/fr-iit/GSMRS。

原文摘要 · Abstract (English)

The movie recommender system typically leverages user feedback to provide personalized recommendations that align with user preferences and increase business revenue. This study investigates the impact of gender stereotypes on such systems through a specific attack scenario. In this scenario, an attacker determines users' gender, a private attribute, by exploiting gender stereotypes about movie preferences and analyzing users' feedback data, which is either publicly available or observed within the system. The study consists of two phases. In the first phase, a user study involving 630 participants identified gender stereotypes associated with movie genres, which often influence viewing choices. In the second phase, four inference algorithms were applied to detect gender stereotypes by combining the findings from the first phase with users' feedback data. Results showed that these algorithms performed more effectively than relying solely on feedback data for gender inference. Additionally, we quantified the extent of gender stereotypes to evaluate their broader impact on digital computational science. The latter part of the study utilized two major movie recommender datasets: MovieLens 1M and Yahoo!Movie. Detailed experimental information is available on our GitHub repository: https://github.com/fr-iit/GSMRS

推荐系统性别偏见隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。