研究图像分类中隐私、公平与效用的权衡,提出新综合指标。
The Impact of Generalization Techniques on the Interplay Among Privacy, Utility, and Fairness in Image Classification
- 引入尖锐感知训练与差分隐私结合,提升模型平衡性。
- 在CIFAR-10上实现81.11%准确率,优于此前79.5%记录。
- 发现泛化技术可能放大偏见,且移除异常样本反而恶化公平性。
本研究探讨机器学习图像分类中公平性、隐私性与效用之间的权衡。近期研究表明,泛化技术可改善隐私与效用的平衡。本文重点分析尖锐感知训练(SAT)及其与差分隐私结合的DP-SAT方法。同时,在含合成与真实世界偏见的数据集上评估私有与非私有模型的公平性。通过成员推断攻击(MIAs)测量隐私风险,并探索剔除高风险样本(即异常值)的影响。此外,提出新指标“调和分数”(harmonic score),统一衡量准确率、隐私与公平性。实证分析显示,在CIFAR-10上以(8, 10⁻⁵)-DP条件下达到81.11%准确率,超越De等(2022)报告的79.5%。实验还发现,训练样本的记忆现象可在过拟合前出现,泛化技术无法保证杜绝记忆。对合成偏见的分析表明,泛化技术会加剧私有与非私有模型的偏见。数据偏见增加导致准确率下降、隐私漏洞扩大及模型偏见上升。在CelebA数据集上的验证表明,真实世界属性失衡下趋势一致。最终实验显示,剔除异常值会降低准确率并进一步放大模型偏见。
原文摘要 · Abstract (English)
This study investigates the trade-offs between fairness, privacy, and utility in image classification using machine learning (ML). Recent research suggests that generalization techniques can improve the balance between privacy and utility. One focus of this work is sharpness-aware training (SAT) and its integration with differential privacy (DP-SAT) to further improve this balance. Additionally, we examine fairness in both private and non-private learning models trained on datasets with synthetic and real-world biases. We also measure the privacy risks involved in these scenarios by performing membership inference attacks (MIAs) and explore the consequences of eliminating high-privacy risk samples, termed outliers. Moreover, we introduce a new metric, named \emph{harmonic score}, which combines accuracy, privacy, and fairness into a single measure. Through empirical analysis using generalization techniques, we achieve an accuracy of 81.11\% under $(8, 10^{-5})$-DP on CIFAR-10, surpassing the 79.5\% reported by De et al. (2022). Moreover, our experiments show that memorization of training samples can begin before the overfitting point, and generalization techniques do not guarantee the prevention of this memorization. Our analysis of synthetic biases shows that generalization techniques can amplify model bias in both private and non-private models. Additionally, our results indicate that increased bias in training data leads to reduced accuracy, greater vulnerability to privacy attacks, and higher model bias. We validate these findings with the CelebA dataset, demonstrating that similar trends persist with real-world attribute imbalances. Finally, our experiments show that removing outlier data decreases accuracy and further amplifies model bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。