arXiv:2505.04038cs.CYcs.LG2025-05被引 11

公平机器学习常忽视歧视形式差异,导致过度泛化。

Identities are not Interchangeable: The Problem of Overgeneralization in Fair Machine Learning

  • 指出公平算法常将种族、性别等歧视混为一谈
  • 强调需关注具体歧视类型以提升公平性
  • 适合关注算法公平与社会正义的研究者

机器学习的核心优势在于泛化能力:同一方法和模型架构可适用于不同领域与情境。然而,这种泛化有时会过度,忽略具体情境的重要性。本文探讨公平机器学习中常将歧视发生的身份维度视为可互换的问题,即对种族主义、性别歧视、能力歧视、年龄歧视等采用相同的衡量与缓解方式。尽管这些歧视存在共性,但非计算机科学领域的研究已揭示其差异。本文由此提出,公平机器学习虽不必对每种歧视都定制方法,但当前对具体歧视形式的关注仍严重不足。强调上下文特异性有助于深化对公平系统的理解,拓展对被忽视伤害的覆盖范围,并在看似无限的群体分析方法中实现有效聚焦。

原文摘要 · Abstract (English)

A key value proposition of machine learning is generalizability: the same methods and model architecture should be able to work across different domains and different contexts. While powerful, this generalization can sometimes go too far, and miss the importance of the specifics. In this work, we look at how fair machine learning has often treated as interchangeable the identity axis along which discrimination occurs. In other words, racism is measured and mitigated the same way as sexism, as ableism, as ageism. Disciplines outside of computer science have pointed out both the similarities and differences between these different forms of oppression, and in this work we draw out the implications for fair machine learning. While certainly not all aspects of fair machine learning need to be tailored to the specific form of oppression, there is a pressing need for greater attention to such specificity than is currently evident. Ultimately, context specificity can deepen our understanding of how to build more fair systems, widen our scope to include currently overlooked harms, and, almost paradoxically, also help to narrow our scope and counter the fear of an infinite number of group-specific methods of analysis.

公平算法歧视识别泛化问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。