arXiv:2411.11101cs.LGcs.AI2024-11被引 6

不同算法在公平性上表现差异大,参数调优比算法本身更重要。

Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML

  • 对比多种偏见缓解算法时,超参数和随机种子影响显著。
  • 多数算法经优化后公平性表现相近,无明显优劣之分。
  • 适合关注算法开发流程对公平性影响的研究者阅读。

随着机器学习中的公平性问题日益受到重视,众多偏见缓解技术被提出,并常通过相互比较以确定最佳方法。这类基准测试通常采用统一评估设置,假设相同环境可保证公平比较。然而,偏见缓解技术对超参数选择、随机种子、特征选取等因素敏感,单一设置下的比较可能不公平地偏向某些算法。本文揭示了多种算法在公平性表现上的显著差异,以及学习管道对公平性评分的影响。研究发现,大多数偏见缓解技术在允许超参数优化的情况下,能达到相近的性能,表明评估参数的选择——而非算法本身——有时会制造出某方法优于另一方法的错觉。我们希望本工作能推动未来研究关注算法生命周期中各项决策如何影响公平性,并为算法选型提供指导。

原文摘要 · Abstract (English)

With fairness concerns gaining significant attention in Machine Learning (ML), several bias mitigation techniques have been proposed, often compared against each other to find the best method. These benchmarking efforts tend to use a common setup for evaluation under the assumption that providing a uniform environment ensures a fair comparison. However, bias mitigation techniques are sensitive to hyperparameter choices, random seeds, feature selection, etc., meaning that comparison on just one setting can unfairly favour certain algorithms. In this work, we show significant variance in fairness achieved by several algorithms and the influence of the learning pipeline on fairness scores. We highlight that most bias mitigation techniques can achieve comparable performance, given the freedom to perform hyperparameter optimization, suggesting that the choice of the evaluation parameters-rather than the mitigation technique itself-can sometimes create the perceived superiority of one method over another. We hope our work encourages future research on how various choices in the lifecycle of developing an algorithm impact fairness, and trends that guide the selection of appropriate algorithms.

公平性偏见缓解超参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。