arXiv:2501.14414cs.LGcs.CR2025-01中稿 · publication in the…被引 3

差分隐私可能放大模型对不同群体的不公平,这篇综述首次系统梳理了原因。

SoK: What Makes Private Learning Unfair?

  • 按机器学习流程分类分析隐私机制如何加剧不公平
  • 数据集和分布特征是导致不公平的核心因素
  • 适合关注隐私与公平平衡的研究者阅读

差分隐私是当前最主流的隐私保护机器学习框架。然而,近期研究表明,强制实施差分隐私不仅会显著降低模型性能,还可能放大不同人群间预测结果的差异。尽管已有大量研究探讨导致该现象的因素,但对其内在机制仍缺乏完整理解。现有文献因公平性定义、差分隐私机制及实验设置不一致,常出现看似矛盾的结果。本文首次全面综述了差分隐私引发不公平影响的关键因素,通过构建分类体系,分析其在机器学习流程中的位置及其因果作用。研究发现,训练数据集特征和底层数据分布是决定不公平效应发生与否的关键,凸显了针对这些因素开展研究的重要性。

原文摘要 · Abstract (English)

Differential privacy has emerged as the most studied framework for privacy-preserving machine learning. However, recent studies show that enforcing differential privacy guarantees can not only significantly degrade the utility of the model, but also amplify existing disparities in its predictive performance across demographic groups. Although there is extensive research on the identification of factors that contribute to this phenomenon, we still lack a complete understanding of the mechanisms through which differential privacy exacerbates disparities. The literature on this problem is muddled by varying definitions of fairness, differential privacy mechanisms, and inconsistent experimental settings, often leading to seemingly contradictory results. This survey provides the first comprehensive overview of the factors that contribute to the disparate effect of training models with differential privacy guarantees. We discuss their impact and analyze their causal role in such a disparate effect. Our analysis is guided by a taxonomy that categorizes these factors by their position within the machine learning pipeline, allowing us to draw conclusions about their interaction and the feasibility of potential mitigation strategies. We find that factors related to the training dataset and the underlying distribution play a decisive role in the occurrence of disparate impact, highlighting the need for research on these factors to address the issue.

差分隐私公平性机器学习综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。