用优化方法选数据准备策略,兼顾公平与性能。
Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
- 提出FATE算法,自动挑选最优数据准备流程。
- 实证显示新方法在公平性和模型性能上优于传统预处理。
- 适合希望快速落地公平性改进的实践者使用。
随着机器学习系统在各行业的广泛应用,解决公平性与偏差问题变得至关重要。尽管许多解决方案聚焦于机器学习中的伦理挑战,但近期研究指出数据本身是偏差的主要来源。预处理技术虽能有效缓解偏差,但可能影响模型性能且集成困难。相比之下,面向公平性的数据准备实践对从业者更熟悉且易于实施,提供了一种更具可访问性的偏差减少途径。本注册报告旨在通过实证评估,在机器学习早期阶段应用最优选择的公平性导向数据准备策略,能否同时提升公平性与性能,甚至超越标准预处理偏差缓解方法。为此,我们将引入FATE——一种用于优化选择‘数据准备’流水线以平衡公平性与性能的算法。利用FATE,我们将分析公平性-性能权衡,并比较其选定的流水线与预处理方法的效果。
原文摘要 · Abstract (English)
As machine learning (ML) systems are increasingly adopted across industries, addressing fairness and bias has become essential. While many solutions focus on ethical challenges in ML, recent studies highlight that data itself is a major source of bias. Pre-processing techniques, which mitigate bias before training, are effective but may impact model performance and pose integration difficulties. In contrast, fairness-aware Data Preparation practices are both familiar to practitioners and easier to implement, providing a more accessible approach to reducing bias. Objective. This registered report proposes an empirical evaluation of how optimally selected fairness-aware practices, applied in early ML lifecycle stages, can enhance both fairness and performance, potentially outperforming standard pre-processing bias mitigation methods. Method. To this end, we will introduce FATE, an optimization technique for selecting 'Data Preparation' pipelines that optimize fairness and performance. Using FATE, we will analyze the fairness-performance trade-off, comparing pipelines selected by FATE with results by pre-processing bias mitigation techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。