arXiv:2503.02968cs.LGcs.CR2025-03被引 1

提出兼顾隐私与公平的合成表格数据生成方法

Privacy-Preserving Fair Synthetic Tabular Data

  • 基于WGAN-GP改进,加入隐私与公平约束
  • 在四个数据集上实现三者平衡,优于现有模型
  • 适合需要合规、无偏数据发布的机构使用

包含敏感信息的表格数据因法律和伦理问题难以共享。合成数据作为替代方案,由机器学习算法生成,试图捕捉真实数据分布。然而,模型存在记忆和偏差风险,依赖训练数据。本文同时解决合成数据的隐私与公平问题,提出PF-WGAN模型,在WGAN-GP基础上引入隐私与公平约束,生成既保护个体隐私又避免群体偏见的数据。在四个不同数据集上与三种先进合成数据生成模型对比,结果显示该模型在实用性、隐私性和公平性之间取得更优平衡。

原文摘要 · Abstract (English)

Sharing of tabular data containing valuable but private information is limited due to legal and ethical issues. Synthetic data could be an alternative solution to this sharing problem, as it is artificially generated by machine learning algorithms and tries to capture the underlying data distribution. However, machine learning models are not free from memorization and may introduce biases, as they rely on training data. Producing synthetic data that preserves privacy and fairness while maintaining utility close to the real data is a challenging task. This research simultaneously addresses both the privacy and fairness aspects of synthetic data, an area not explored by other studies. In this work, we present PF-WGAN, a privacy-preserving, fair synthetic tabular data generator based on the WGAN-GP model. We have modified the original WGAN-GP by adding privacy and fairness constraints forcing it to produce privacy-preserving fair data. This approach will enable the publication of datasets that protect individual's privacy and remain unbiased toward any particular group. We compared the results with three state-of-the-art synthetic data generator models in terms of utility, privacy, and fairness across four different datasets. We found that the proposed model exhibits a more balanced trade-off among utility, privacy, and fairness.

合成数据隐私保护公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。