arXiv:2609.03180cs.LG2026-09

跨生成器家族实现可移植的因果公平性,保障数据隐私下仍保持高公平性与精度。

Portable Causal Fairness Across Synthetic Data Generator Families

  • 基于因果图边裁剪机制,在九类生成器中统一实现三类公平性定义。
  • 公平性提升显著,仅牺牲0.07~0.15 AUC,且隐私保护不削弱公平性。
  • 适用于需公平性保障的数据发布场景,尤其适合监管与统计机构使用。

当统计机构或监管方以合成数据替代敏感记录时,可通过设计生成器消除不公平路径。DECAF在单一非私有GAN上验证了三种公平性定义对应因果图中的边裁剪。本研究将该机制扩展至三个无关生成器家族(基于边际、GAN、扩散模型,含差分隐私变体),在成人和COMPAS数据集上进行2520次配对实验,覆盖三种形式化隐私保证水平。结果表明,该机制在所有生成器中均有效迁移;新提出的因果扩散模型骨架在测试的所有家族中实现最优公平性,同时保有接近边际级的保真度。边裁剪对保真度影响极小,下游分类器平均损失0.07~0.15 AUC,添加隐私保护不会降低公平性。

原文摘要 · Abstract (English)

When a statistical agency or regulator releases synthetic data in place of sensitive records, it chooses the generator that produces the table, and can shape that generator so unfair pathways are absent. DECAF made this concrete on one non-private GAN: three fairness definitions become three sets of edge cuts on the generator's causal graph. Whether the mechanism belongs to DECAF, or to causal factorisation itself, was untested. We port all three definitions to nine generators from three unrelated families (marginals-based, GAN, and diffusion, each with differentially private variants), across three levels of formal privacy guarantee, over 2,520 matched-pair runs on Adult and COMPAS datasets. The mechanism transfers everywhere, and our new causal diffusion backbone yields the fairest release of any family we tested, at fidelity close to the marginals tier. Applying the cut barely moves fidelity, only costs a downstream classifier about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees don't make the data less fair.

因果公平合成数据差分隐私生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。