构建可调控偏见的合成数据集,揭示模型如何继承并放大数据中的不公平。
Simulating Biases for Interpretable Fairness in Offline and Online Classifiers
- 用基于代理的模型模拟贷款审批过程,生成含可控偏见的合成数据
- 实验表明,预处理与训练中干预可显著降低模型对群体的不公平性
- 提出第二阶Shapley值解释技术,可视化公平性修正如何改变特征使用方式
预测模型常因训练数据中的固有偏见而加剧决策不公。为评估此类问题,需系统性生成含特定偏见的数据集以训练分类器,并分析偏见传播机制。本文提出一种基于代理的模型(ABM),模拟跨两个群体的贷款申请流程,显式建模多种系统性偏见,生成合成数据集。在此基础上,训练分类器并评估其预测结果,揭示偏见如何导致不公平。主要贡献包括:一个支持可控偏见注入的合成数据生成框架;以及一种新型可解释性方法,利用第二阶Shapley值展示公平性缓解措施如何影响分类器对特征的依赖模式。在离线与在线学习场景下进行实验,分别在预处理与训练阶段应用缓解策略,验证了其有效性。
原文摘要 · Abstract (English)
Predictive models often reinforce biases which were originally embedded in their training data, through skewed decisions. In such cases, mitigation methods are critical to ensure that, regardless of the prevailing disparities, model outcomes are adjusted to be fair. To assess this, datasets could be systematically generated with specific biases, to train machine learning classifiers. Then, predictive outcomes could aid in the understanding of this bias embedding process. Hence, an agent-based model (ABM), depicting a loan application process that represents various systemic biases across two demographic groups, was developed to produce synthetic datasets. Then, by applying classifiers trained on them to predict loan outcomes, we can assess how biased data leads to unfairness. This highlights a main contribution of this work: a framework for synthetic dataset generation with controllable bias injection. We also contribute with a novel explainability technique, which shows how mitigations affect the way classifiers leverage data features, via second-order Shapley values. In experiments, both offline and online learning approaches are employed. Mitigations are applied at different stages of the modelling pipeline, such as during pre-processing and in-processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。