arXiv:2501.01785cs.LGcs.AI2025-01被引 24

合成数据能兼顾公平与隐私吗?研究发现组合使用特定生成器和公平性算法效果最佳。

Can Synthetic Data be Fair and Private? A Comparative Study of Synthetic Data Generation and Fairness Algorithms

  • 选用DECAF生成器并结合预处理公平算法,实现隐私与公平的最优平衡。
  • 在合成数据上应用公平算法,公平性提升优于真实数据上的效果。
  • 尽管准确性略有下降,但该方法为教育数据分析提供更安全可靠的解决方案。

机器学习在学习分析(LA)中的广泛应用引发了算法公平性和隐私保护的重大关切。合成数据作为一种双重工具,既能增强隐私保护,又能提升模型公平性。然而,先前研究指出公平性与隐私之间存在负相关关系,难以同时优化。本研究探讨了不同合成数据生成器在平衡隐私与公平方面的表现,并检验了通常用于真实数据的预处理公平算法在合成数据上的有效性。结果表明,去偏因果公平(DECAF)算法在隐私与公平的权衡中表现最佳。然而,该方法在实用性上有所损失,表现为预测准确率下降。值得注意的是,在合成数据上应用预处理公平算法,其公平性提升效果甚至超过在真实数据上的表现。研究结果表明,将合成数据生成与公平性预处理相结合,是一种有前景的构建更公平学习分析模型的方法。

原文摘要 · Abstract (English)

The increasing use of machine learning in learning analytics (LA) has raised significant concerns around algorithmic fairness and privacy. Synthetic data has emerged as a dual-purpose tool, enhancing privacy and improving fairness in LA models. However, prior research suggests an inverse relationship between fairness and privacy, making it challenging to optimize both. This study investigates which synthetic data generators can best balance privacy and fairness, and whether pre-processing fairness algorithms, typically applied to real datasets, are effective on synthetic data. Our results highlight that the DEbiasing CAusal Fairness (DECAF) algorithm achieves the best balance between privacy and fairness. However, DECAF suffers in utility, as reflected in its predictive accuracy. Notably, we found that applying pre-processing fairness algorithms to synthetic data improves fairness even more than when applied to real data. These findings suggest that combining synthetic data generation with fairness pre-processing offers a promising approach to creating fairer LA models.

合成数据公平性隐私保护学习分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。