arXiv:2607.07471cs.LGcs.AI2026-07中稿 · PETS 2026

首次系统评估隐私保护数据下的公平性干预效果,发现事后修正最稳定有效。

Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data

论文配图:Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
图 1 · 摘自论文原文
  • 在差分隐私合成数据上测试三种公平干预策略,比较其效果与稳定性。
  • 事后处理方法在不同隐私预算下均能提升公平性,且保持较好数据效用。
  • 研究揭示隐私与公平的权衡关系,适合关注数据伦理与模型公平的研究者。

机器学习模型在高风险领域应用日益广泛,引发对隐私与公平的双重担忧。差分隐私(DP)已成为隐私保护数据处理的标准,而公平性干预旨在缓解对少数群体的歧视。然而,两者存在冲突:DP常加剧群体间差异,现有公平机制在DP约束下是否仍有效尚不明确。本文首次系统评估公平干预在差分隐私合成表格数据上的表现,聚焦于当前最优的基于边际的合成机制AIM(Cormode et al. 2025)。我们在四个数据集上,使用多种群体公平性度量,评估三类干预策略(预处理、内处理、后处理),覆盖广泛的隐私预算。对比四种流程:(基线)原始数据训练;(仅DP)DP合成数据训练;(仅公平)原始数据加公平机制;(DP+公平)结合两者。结果表明,尽管DP会降低效用与公平性,但引入公平干预可部分恢复公平性。其中,后处理方法在不同隐私预算和合成器下表现更稳定,显著提升公平性的同时保持较高效用。代码、数据及实验资产均已开源,支持未来研究。

原文摘要 · Abstract (English)

Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups. However, these objectives can conflict: DP often amplifies disparities across demographic groups, and little is known about whether established fairness interventions remain effective under DP constraints. In this work, we present, to our knowledge, the first systematic evaluation of fairness interventions on differentially private synthetic tabular data. Our benchmark centers on the Adaptive Iterative Mechanism (AIM), identified as the state-of-the-art marginal-based DP synthesizer (Cormode et al. 2025). We thus evaluate fairness interventions across four datasets, multiple group fairness metrics, and three categories of mitigation strategies (pre-processing, in-processing, and post-processing) under a wide range of privacy budgets. We compare four pipeline configurations: (Baseline) training on original data; (DP-only) training on DP synthetic data; (Fair-only) applying fairness mechanisms on original data; and (DP+Fair) combining fairness mechanisms with DP synthetic data. Our results demonstrate that while DP alone can degrade both utility and fairness, applying fairness interventions can partially restore equitable outcomes. Among them, post-processing methods tend to provide more stable fairness-utility trade-offs across privacy budgets and synthesizers, achieving strong fairness improvements while preserving competitive utility relative to other intervention stages. We release all code, data, and experimental artifacts in an open-source repository to ensure full reproducibility and to support future research on the privacy-fairness-utility trade-off.

差分隐私公平性合成数据机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。