用生成模型合成数据,缓解分类与群体不平衡问题。
Synthetic Tabular Data Generation for Class Imbalance and Fairness: A Comparative Study
- 采用先进生成模型合成表格数据,替代传统过采样方法。
- 在4个真实数据集上验证,生成数据显著提升公平性与模型性能。
- 适合关注数据偏见、公平性的机器学习研究者参考。
由于机器学习(ML)模型具有数据驱动特性,易继承数据中的偏见,尤其在分类任务中,类别不平衡和受保护属性(如性别或种族)的群体不平衡普遍存在。这两类不平衡常同时出现,但现有方法对此关注有限。尽管多数方法使用插值等过采样技术缓解问题,近年来合成表格数据生成技术展现出潜力,却尚未充分探索其在该场景的应用。本文对最先进的合成表格数据生成模型及多种采样策略进行对比分析,旨在解决类别与群体双重不平衡问题。在四个真实数据集上的实验表明,生成模型能有效缓解偏见,为该方向提供新思路。
原文摘要 · Abstract (English)
Due to their data-driven nature, Machine Learning (ML) models are susceptible to bias inherited from data, especially in classification problems where class and group imbalances are prevalent. Class imbalance (in the classification target) and group imbalance (in protected attributes like sex or race) can undermine both ML utility and fairness. Although class and group imbalances commonly coincide in real-world tabular datasets, limited methods address this scenario. While most methods use oversampling techniques, like interpolation, to mitigate imbalances, recent advancements in synthetic tabular data generation offer promise but have not been adequately explored for this purpose. To this end, this paper conducts a comparative analysis to address class and group imbalances using state-of-the-art models for synthetic tabular data generation and various sampling strategies. Experimental results on four datasets, demonstrate the effectiveness of generative models for bias mitigation, creating opportunities for further exploration in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。