提出合成数据生成中的公平性新标准,强调真实与合成分布一致才算公平。
Disparate Impact in Synthetic Data Generation
- 以真实分布一致性为公平标准,突破传统偏差修正思路。
- 发现模型表达能力、群体比例、差分隐私等因素导致不同群体误差不均。
- 建议分组训练合成模型,提升整体性能与公平性,适合数据隐私与公平研究者。
我们重新审视合成数据生成(SDG)中的歧视性影响公平性概念,该概念评估生成记录在敏感群体间的效用是否一致。现有方法通常通过纠正观测分布中的偏见来实现公平,将SDG重新定义为学习非真实数据分布的生成模型。而本文主张,当合成分布与真实分布相同时,才真正实现无歧视性影响。我们揭示了当前SDG方法难以达成此目标的原因,包括模型表达能力不足、分布复杂度高、群体比例导致的采样误差,以及差分隐私机制引发的估计偏差,且这些误差在不同群体间可能不均等。我们在人工和真实数据上展示了多种基于概率图模型的SDG方法存在的歧视性影响。此外,我们提出分组学习合成模型的策略,实证表明该方法可在多个场景中同时提升整体效用与群体间公平性。
原文摘要 · Abstract (English)
We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data. By contrast, non-disparate impact is notably achieved when the synthetic and real distributions are the same. We expose reasons why SDG may fail to reach that solution and discuss why approximation and estimation errors occur and can be disparate across groups. We notably look into the expressive power of SDG methods relative to distribution complexity, sampling errors due to group proportions, and estimation errors induced by differential privacy mechanisms. We illustrate cases of disparate impact on both artificial and real-world data, focusing on SDG methods that rely on probabilistic graphical models. We also introduce a strategy of learning group-wise SDG models and illustrate how it can improve both the overall utility and its parity in many settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。