arXiv:2410.01423cs.LGcs.AI2024-10被引 1

无需原始数据,生成公平高质量合成数据

Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data Free Distillation

  • 在隐空间中用无数据蒸馏生成公平表示
  • 合成数据在公平性、实用性、质量上均提升5%-12%
  • 适合隐私保护场景下的数据增强与公平性改进

本文提出Fair4Free,一种基于隐空间无数据蒸馏的生成模型,可在数据私密或不可访问时生成公平的合成数据。方法先训练教师模型以生成公平表征,再将知识蒸馏至参数更小的学生模型,整个蒸馏过程不依赖原始训练数据。蒸馏完成后,使用学生模型生成公平合成样本。大量实验表明,该方法在表格和图像数据集上,合成样本在公平性、实用性及合成质量三个指标上均优于现有最优模型,分别提升5%、8%和12%。

原文摘要 · Abstract (English)

This work presents Fair4Free, a novel generative model to generate synthetic fair data using data-free distillation in the latent space. Fair4Free can work on the situation when the data is private or inaccessible. In our approach, we first train a teacher model to create fair representation and then distil the knowledge to a student model (using a smaller architecture). The process of distilling the student model is data-free, i.e. the student model does not have access to the training dataset while distilling. After the distillation, we use the distilled model to generate fair synthetic samples. Our extensive experiments show that our synthetic samples outperform state-of-the-art models in all three criteria (fairness, utility and synthetic quality) with a performance increase of 5% for fairness, 8% for utility and 12% in synthetic quality for both tabular and image datasets.

合成数据公平性无数据蒸馏隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。