arXiv:2505.07985cs.LGcs.AI2025-05被引 1

匿名化技术会严重损害群体公平性,却可能提升个体公平性。

Fair Play for Individuals, Foul Play for Groups? Auditing Anonymization's Impact on ML Fairness

  • 系统评估k-匿名、ℓ-多样性等技术对机器学习公平性的影响。
  • 匿名化使群体公平性指标下降最高达4倍,但个体相似性公平性提升。
  • 揭示隐私、公平与模型效用间的权衡,适合关注负责任AI的研究者。

机器学习依赖大量训练数据,其中常包含敏感信息,引发严重隐私担忧。匿名化技术(如k-匿名、ℓ-多样性、t- closeness)通过泛化或抑制数据特征来降低个体识别风险,被视为实用隐私保护方案。尽管已有研究指出隐私技术会影响不同子群体的预测结果,但其对机器学习公平性的具体影响仍不明确。本文系统审计了多种匿名化技术对个体与群体公平性的影响。定量分析显示,匿名化可使群体公平性指标恶化最高达四倍;而基于相似性的个体公平性则因输入同质性增强而提升。通过在多样隐私设置和数据分布下分析不同匿名化程度,本研究揭示了隐私、公平与模型效用之间的关键权衡,为负责任的人工智能开发提供可操作指导。代码已公开于:https://github.com/hharcolezi/anonymity-impact-fairness。

原文摘要 · Abstract (English)

Machine learning (ML) algorithms are heavily based on the availability of training data, which, depending on the domain, often includes sensitive information about data providers. This raises critical privacy concerns. Anonymization techniques have emerged as a practical solution to address these issues by generalizing features or suppressing data to make it more difficult to accurately identify individuals. Although recent studies have shown that privacy-enhancing technologies can influence ML predictions across different subgroups, thus affecting fair decision-making, the specific effects of anonymization techniques, such as $k$-anonymity, $\ell$-diversity, and $t$-closeness, on ML fairness remain largely unexplored. In this work, we systematically audit the impact of anonymization techniques on ML fairness, evaluating both individual and group fairness. Our quantitative study reveals that anonymization can degrade group fairness metrics by up to fourfold. Conversely, similarity-based individual fairness metrics tend to improve under stronger anonymization, largely as a result of increased input homogeneity. By analyzing varying levels of anonymization across diverse privacy settings and data distributions, this study provides critical insights into the trade-offs between privacy, fairness, and utility, offering actionable guidelines for responsible AI development. Our code is publicly available at: https://github.com/hharcolezi/anonymity-impact-fairness.

隐私保护机器学习公平性匿名化负责任AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。