arXiv:2409.12428cs.LGcs.AI2024-09被引 4

研究数据分布漂移如何影响公平算法效果,发现看似公平的模型可能因数据变化而变不公平。

Is it Still Fair? A Comparative Evaluation of Fairness Algorithms through the Lens of Covariate Drift

  • 对比11种算法在5个数据集上的表现,考察数据漂移对公平性的影响。
  • 发现数据漂移会严重恶化所谓'公平模型'的公平性,且与漂移方向无关。
  • 提醒从业者关注数据动态变化,否则公平性措施可能失效。

近年来,机器学习应用迅速发展,但其潜在的歧视性行为引发广泛关注。公平性成为机器学习的重要研究方向,已有多种公平性度量和算法被提出。然而,数据分布漂移(即数据模式自然变化)对公平性算法和度量的影响仍鲜受关注。本文系统评估了4种公平性无感知基线算法和7种公平性感知算法,在包括公开与私有数据在内的5个数据集上,使用3项预测性能指标和10项公平性指标进行测试。结果表明:(1)数据分布漂移并非偶然现象,常导致所谓‘公平模型’的公平性显著下降;(2)数据漂移的大小与方向与不公平性的变化无相关性;(3)公平算法的选择与训练受数据漂移影响,但这一问题在现有文献中普遍被忽略。基于研究发现,本文提出了若干对实践者和决策者的政策启示。

原文摘要 · Abstract (English)

Over the last few decades, machine learning (ML) applications have grown exponentially, yielding several benefits to society. However, these benefits are tempered with concerns of discriminatory behaviours exhibited by ML models. In this regard, fairness in machine learning has emerged as a priority research area. Consequently, several fairness metrics and algorithms have been developed to mitigate against discriminatory behaviours that ML models may possess. Yet still, very little attention has been paid to the problem of naturally occurring changes in data patterns (\textit{aka} data distributional drift), and its impact on fairness algorithms and metrics. In this work, we study this problem comprehensively by analyzing 4 fairness-unaware baseline algorithms and 7 fairness-aware algorithms, carefully curated to cover the breadth of its typology, across 5 datasets including public and proprietary data, and evaluated them using 3 predictive performance and 10 fairness metrics. In doing so, we show that (1) data distributional drift is not a trivial occurrence, and in several cases can lead to serious deterioration of fairness in so-called fair models; (2) contrary to some existing literature, the size and direction of data distributional drift is not correlated to the resulting size and direction of unfairness; and (3) choice of, and training of fairness algorithms is impacted by the effect of data distributional drift which is largely ignored in the literature. Emanating from our findings, we synthesize several policy implications of data distributional drift on fairness algorithms that can be very relevant to stakeholders and practitioners.

公平性数据漂移算法评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。