arXiv:2504.05923cs.LGcs.AI2025-04

用分类复杂度差异提前发现模型不公平性

Uncovering Fairness through Data Complexity as an Early Indicator

  • 通过合成数据测试不同群体的分类复杂度差异
  • 复杂度不均与公平性指标显著相关,可预判偏见风险
  • 为算法开发者提供早期干预的数据依据

公平性是机器学习应用中的关键问题。目前尚无研究探讨特权群体与非特权群体在分类复杂度上的差异如何影响模型公平性,而这可能成为潜在不公平性的早期指示信号。本文通过设计包含历史偏差、测量偏差和表征偏差等多种偏见的合成数据集,评估各类复杂度指标在群体间的差异,并与群体公平性指标进行关联分析。利用关联规则挖掘技术,识别出群体间复杂度差异与公平性结果之间的模式,提出以数据为中心的早期公平性预警指标。研究结果在真实场景中得到验证,表明量化群体分类复杂度有助于提前发现潜在的公平性挑战,帮助从业者主动应对分类任务中的偏见问题。

原文摘要 · Abstract (English)

Fairness constitutes a concern within machine learning (ML) applications. Currently, there is no study on how disparities in classification complexity between privileged and unprivileged groups could influence the fairness of solutions, which serves as a preliminary indicator of potential unfairness. In this work, we investigate this gap, specifically, we focus on synthetic datasets designed to capture a variety of biases ranging from historical bias to measurement and representational bias to evaluate how various complexity metrics differences correlate with group fairness metrics. We then apply association rule mining to identify patterns that link disproportionate complexity differences between groups with fairness-related outcomes, offering data-centric indicators to guide bias mitigation. Our findings are also validated by their application in real-world problems, providing evidence that quantifying group-wise classification complexity can uncover early indicators of potential fairness challenges. This investigation helps practitioners to proactively address bias in classification tasks.

公平性数据复杂度偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。