arXiv:2410.15470cs.LGcs.AI2024-10被引 11

用扩散模型生成表格数据,提升AI分类公平性。

Data Augmentation via Diffusion Model to Enhance AI Fairness

  • 基于扩散模型生成合成表格数据,适配多种特征类型。
  • 在五种模型上验证,合成数据显著提升二分类公平性。
  • 适合关注数据偏见与公平性的机器学习研究者。

AI公平性旨在通过确保系统决策真正反映用户最佳利益来提升AI的透明度与可解释性。数据增强通过从现有数据集中生成合成数据,成为缓解数据稀缺的重要手段。特别是扩散模型在计算机视觉等领域展现出强大的合成数据生成能力。本文探索了扩散模型生成合成表格数据以提升AI公平性的潜力。采用可适配任意表格数据集且支持多种特征类型的表格式去噪扩散概率模型(Tab-DDPM),并测试不同数量的生成数据用于数据增强。同时结合AIF360数据集的样本重加权策略进一步优化公平性。使用决策树(DT)、高斯朴素贝叶斯(GNB)、K近邻(KNN)、逻辑回归(LR)和随机森林(RF)五种经典机器学习模型验证方法有效性。实验结果表明,由Tab-DDPM生成的合成数据能有效提升二分类任务中的公平性表现。

原文摘要 · Abstract (English)

AI fairness seeks to improve the transparency and explainability of AI systems by ensuring that their outcomes genuinely reflect the best interests of users. Data augmentation, which involves generating synthetic data from existing datasets, has gained significant attention as a solution to data scarcity. In particular, diffusion models have become a powerful technique for generating synthetic data, especially in fields like computer vision. This paper explores the potential of diffusion models to generate synthetic tabular data to improve AI fairness. The Tabular Denoising Diffusion Probabilistic Model (Tab-DDPM), a diffusion model adaptable to any tabular dataset and capable of handling various feature types, was utilized with different amounts of generated data for data augmentation. Additionally, reweighting samples from AIF360 was employed to further enhance AI fairness. Five traditional machine learning models-Decision Tree (DT), Gaussian Naive Bayes (GNB), K-Nearest Neighbors (KNN), Logistic Regression (LR), and Random Forest (RF)-were used to validate the proposed approach. Experimental results demonstrate that the synthetic data generated by Tab-DDPM improves fairness in binary classification.

数据增强扩散模型公平性表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。