arXiv:2410.07182cs.IRcs.CY2024-10被引 3

研究推荐系统中数据最小化与公平性的权衡,发现主动学习会损害公平性。

The trade-off between data minimization and fairness in collaborative filtering

  • 用主动学习实现数据最小化,兼顾高精度与少数据采集
  • 多数主动学习策略使推荐系统公平性下降,但准确性保持较高
  • 为合规欧盟隐私法规的推荐系统设计提供关键权衡参考

通用数据保护条例(GDPR)要求同时遵守公平性、准确性与数据最小化等原则,但在实际应用中这些原则可能存在冲突。本文研究推荐系统中数据最小化与公平性之间的权衡关系。通过主动学习(AL)实现数据最小化,因其可在保证高精度的同时减少数据收集量。在两个公开数据集上对比了个性化与非个性化主动学习策略,结果表明:几乎所有策略均对公平性产生负面影响,尽管准确率保持较高。当前关于数据最小化与公平性权衡、主动学习作为实现手段的利弊及其对公平性的影响研究仍极为有限。本研究填补了这一空白,为构建符合GDPR要求的推荐系统提供了重要实践指导。

原文摘要 · Abstract (English)

General Data Protection Regulations (GDPR) aim to safeguard individuals' personal information from harm. While full compliance is mandatory in the European Union and the California Privacy Rights Act (CPRA), it is not in other places. GDPR requires simultaneous compliance with all the principles such as fairness, accuracy, and data minimization. However, it overlooks the potential contradictions within its principles. This matter gets even more complex when compliance is required from decision-making systems. Therefore, it is essential to investigate the feasibility of simultaneously achieving the goals of GDPR and machine learning, and the potential tradeoffs that might be forced upon us. This paper studies the relationship between the principles of data minimization and fairness in recommender systems. We operationalize data minimization via active learning (AL) because, unlike many other methods, it can preserve a high accuracy while allowing for strategic data collection, hence minimizing the amount of data collection. We have implemented several active learning strategies (personalized and non-personalized) and conducted a comparative analysis focusing on accuracy and fairness on two publicly available datasets. The results demonstrate that different AL strategies may have different impacts on the accuracy of recommender systems with nearly all strategies negatively impacting fairness. There has been no to very limited work on the trade-off between data minimization and fairness, the pros and cons of active learning methods as tools for implementing data minimization, and the potential impacts of AL on fairness. By exploring these critical aspects, we offer valuable insights for developing recommender systems that are GDPR compliant.

推荐系统数据最小化公平性主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。