arXiv:2410.18749cs.CLcs.AI2024-10被引 2

差分隐私会加剧大模型对少数群体的偏见,影响其判断能力。

Does Differential Privacy Impact Bias in Pretrained NLP Models?

  • 通过实验分析差分隐私对预训练模型偏见的影响
  • DP使模型更难区分少数群体的正负样本,导致偏见上升
  • 偏见程度受隐私级别和数据分布共同影响,适合关注公平性的研究者

差分隐私(DP)在微调预训练大语言模型时被用于限制训练样本泄露。尽管多数DP研究聚焦于提升隐私-效用权衡,但已有研究发现其可能对代表性不足群体造成不公或偏见。本文通过实证分析揭示了DP对大模型偏见的影响:差分隐私训练会使模型在基于AUC的偏见度量下对受保护群体的偏见加剧。这表明,DP使模型更难以区分受保护群体与其余群体中的正负样本。结果还显示,DP对偏见的影响不仅取决于隐私保护水平,也受数据集底层分布的影响。

原文摘要 · Abstract (English)

Differential privacy (DP) is applied when fine-tuning pre-trained large language models (LLMs) to limit leakage of training examples. While most DP research has focused on improving a model's privacy-utility tradeoff, some find that DP can be unfair to or biased against underrepresented groups. In this work, we show the impact of DP on bias in LLMs through empirical analysis. Differentially private training can increase the model bias against protected groups w.r.t AUC-based bias metrics. DP makes it more difficult for the model to differentiate between the positive and negative examples from the protected groups and other groups in the rest of the population. Our results also show that the impact of DP on bias is not only affected by the privacy protection level but also the underlying distribution of the dataset.

差分隐私模型偏见大模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。