通过优化数据预处理,实现非二元属性的公平与隐私保护。
Towards Fairness and Privacy: A Novel Data Pre-processing Optimization Framework for Non-binary Protected Attributes
- 构建组合优化框架,用遗传算法减少数据偏见。
- 在真实数据集上验证,生成的公平数据集显著降低歧视度。
- 支持数据删除、合成数据添加等灵活场景,适合隐私敏感应用。
AI不公平结果常源于数据偏见。本文提出一种针对非二元受保护属性的数据预处理框架,通过构建组合优化问题,利用遗传算法等启发式方法实现公平性目标。该框架旨在寻找最小化特定歧视度量的数据子集,支持数据移除、合成数据添加或仅使用合成数据等多种应用场景。尤其当仅使用合成数据时,可同时增强隐私保护能力。在全面评估中,实验表明遗传算法能有效生成比原始数据更公平的数据集。相比已有工作,本框架具备高度灵活性:对度量和任务无关,适用于二元与非二元受保护属性,并具有高效运行时间。
原文摘要 · Abstract (English)
The reason behind the unfair outcomes of AI is often rooted in biased datasets. Therefore, this work presents a framework for addressing fairness by debiasing datasets containing a (non-)binary protected attribute. The framework proposes a combinatorial optimization problem where heuristics such as genetic algorithms can be used to solve for the stated fairness objectives. The framework addresses this by finding a data subset that minimizes a certain discrimination measure. Depending on a user-defined setting, the framework enables different use cases, such as data removal, the addition of synthetic data, or exclusive use of synthetic data. The exclusive use of synthetic data in particular enhances the framework's ability to preserve privacy while optimizing for fairness. In a comprehensive evaluation, we demonstrate that under our framework, genetic algorithms can effectively yield fairer datasets compared to the original data. In contrast to prior work, the framework exhibits a high degree of flexibility as it is metric- and task-agnostic, can be applied to both binary or non-binary protected attributes, and demonstrates efficient runtime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。