研究标签和选择偏差如何影响模型评估与公平性,发现无偏测试集下公平与准确不冲突。
No evaluation without fair representation : Impact of label and selection bias on the evaluation, performance and mitigation of classification models
- 在真实数据中控制引入偏差,独立分析各类偏差影响
- 无偏测试集下公平与准确、个体与群体公平无权衡关系
- 不同偏差类型显著影响公平性缓解方法的效果
机器学习数据集中可能存在多种偏差,如选择偏差和标签偏差。尽管这些偏差对公平机器学习的重要方面有影响,但其差异影响仍研究不足。本文通过实证分析标签偏差及若干子类型选择偏差对分类模型评估、性能以及公平性缓解方法有效性的影响。我们提出一个可建模公平世界及其偏差版本的框架,通过在低歧视真实数据集中引入可控偏差,独立评估每种偏差的影响,获得比传统使用部分偏差数据作为测试集更具代表性的模型与缓解方法评估结果。结果显示,偏差对模型性能的影响受多种因素调节;当在无偏测试集上评估时,公平性与准确性之间不存在权衡,个体公平与群体公平亦无冲突。此外,缓解方法的性能取决于数据中实际存在的偏差类型。研究呼吁未来需发展更精确的模型与干预评估方法,并深入理解其他偏差类型、复杂偏差组合及数据特征等对缓解效率的影响。
原文摘要 · Abstract (English)
Bias can be introduced in diverse ways in machine learning datasets, for example via selection or label bias. Although these bias types in themselves have an influence on important aspects of fair machine learning, their different impact has been understudied. In this work, we empirically analyze the effect of label bias and several subtypes of selection bias on the evaluation of classification models, on their performance, and on the effectiveness of bias mitigation methods. We also introduce a biasing and evaluation framework that allows to model fair worlds and their biased counterparts through the introduction of controlled bias in real-life datasets with low discrimination. Using our framework, we empirically analyze the impact of each bias type independently, while obtaining a more representative evaluation of models and mitigation methods than with the traditional use of a subset of biased data as test set. Our results highlight different factors that influence how impactful bias is on model performance. They also show an absence of trade-off between fairness and accuracy, and between individual and group fairness, when models are evaluated on a test set that does not exhibit unwanted bias. They furthermore indicate that the performance of bias mitigation methods is influenced by the type of bias present in the data. Our findings call for future work to develop more accurate evaluations of prediction models and fairness interventions, but also to better understand other types of bias, more complex scenarios involving the combination of different bias types, and other factors that impact the efficiency of the mitigation methods, such as dataset characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。