在隐私保护下生成公平的表格数据,兼顾差分隐私与群体公平性。
Achieving Hilbert-Schmidt Independence Under Rényi Differential Privacy for Fair and Private Data Generation
- 基于Transformer的变分自编码器结合潜在扩散,实现隐私保护下的数据生成。
- 在多种下游任务中,公平性显著提升且满足瑞尼差分隐私约束。
- 适合医疗等敏感领域使用,尤其关注无任务依赖的公平数据生成。
随着GDPR、HIPAA及AI法案等隐私法规和责任框架的推行,真实世界数据的伦理使用面临更多限制。合成数据生成成为风险可控的数据共享与模型开发的有力方案,尤其适用于医疗等敏感领域的表格数据。为同时解决隐私与公平性问题,我们提出FLIP(公平潜空间干预,带隐私保障),一种基于Transformer的变分自编码器,融合潜在扩散机制以生成异构表格数据。不同于传统公平性导向的数据生成需依赖特定下游任务,本方法采用任务无关设定,适用性更广。为保障隐私,训练过程中引入瑞尼差分隐私(RDP)约束;在输入空间,通过适配组别噪声水平的平衡采样策略实现公平性控制。在潜空间中,利用中心化核对齐(CKA)——一种扩展希尔伯特-施密特独立性准则(HSIC)的相似性度量——对齐受保护群体间的神经元激活模式,促进潜在表示与保护特征的统计独立。实验表明,FLIP在任务无关公平性及多样下游任务中均有效提升了公平性表现,同时满足差分隐私要求。
原文摘要 · Abstract (English)
As privacy regulations such as the GDPR and HIPAA and responsibility frameworks for artificial intelligence such as the AI Act gain traction, the ethical and responsible use of real-world data faces increasing constraints. Synthetic data generation has emerged as a promising solution to risk-aware data sharing and model development, particularly for tabular datasets that are foundational to sensitive domains such as healthcare. To address both privacy and fairness concerns in this setting, we propose FLIP (Fair Latent Intervention under Privacy guarantees), a transformer-based variational autoencoder augmented with latent diffusion to generate heterogeneous tabular data. Unlike the typical setup in fairness-aware data generation, we assume a task-agnostic setup, not reliant on a fixed, defined downstream task, thus offering broader applicability. To ensure privacy, FLIP employs Rényi differential privacy (RDP) constraints during training and addresses fairness in the input space with RDP-compatible balanced sampling that accounts for group-specific noise levels across multiple sampling rates. In the latent space, we promote fairness by aligning neuron activation patterns across protected groups using Centered Kernel Alignment (CKA), a similarity measure extending the Hilbert-Schmidt Independence Criterion (HSIC). This alignment encourages statistical independence between latent representations and the protected feature. Empirical results demonstrate that FLIP effectively provides significant fairness improvements for task-agnostic fairness and across diverse downstream tasks under differential privacy constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。