同时提升数据隐私与公平性的合成数据生成方法
SAFES: Sequential Privacy and Fairness Enhancing Data Synthesis for Responsible AI
- 分步融合差分隐私合成与公平性预处理
- 合理隐私损失下公平性显著提升,效用损失小
- 适用于需兼顾隐私与公平的AI场景
随着数据驱动和基于AI的决策在各领域广泛应用,数据隐私与决策公平性至关重要。尽管差分隐私(DP)提供了可靠的隐私保障框架,且已有方法可提升公平性,但多数先前工作将两者分开处理。即使存在同时考虑隐私与公平的方法,也通常局限于特定学习任务,泛化能力有限。为此,我们提出SAFES——一种顺序增强隐私与公平的数据合成流程,将差分隐私数据合成与公平性感知的数据预处理步骤依次结合。SAFES使用户能够在隐私-公平性-效用之间灵活权衡。我们使用多种DP合成器与公平性预处理方法验证SAFES,并在多个真实数据集上进行广泛实验,评估其生成的合成数据在隐私-公平性-效用之间的平衡表现。实证结果表明,在合理隐私损失下,SAFES生成的合成数据可实现显著改善的公平性指标,同时保持较低的效用损失。
原文摘要 · Abstract (English)
As data-driven and AI-based decision making gains widespread adoption across disciplines, it is crucial that both data privacy and decision fairness are appropriately addressed. Although differential privacy (DP) provides a robust framework for guaranteeing privacy and methods are available to improve fairness, most prior work treats the two concerns separately. Even though there are existing approaches that consider privacy and fairness simultaneously, they typically focus on a single specific learning task, limiting their generalizability. In response, we introduce SAFES, a Sequential PrivAcy and Fairness Enhancing data Synthesis procedure that sequentially combines DP data synthesis with a fairness-aware data preprocessing step. SAFES allows users flexibility in navigating the privacy-fairness-utility trade-offs. We illustrate SAFES with different DP synthesizers and fairness-aware data preprocessing methods and run extensive experiments on multiple real datasets to examine the privacy-fairness-utility trade-offs of synthetic data generated by SAFES. Empirical evaluations demonstrate that for reasonable privacy loss, SAFES-generated synthetic data can achieve significantly improved fairness metrics with relatively low utility loss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。