arXiv:2412.16769cs.LG2024-12被引 3

校准不等于公平,因为个体属于多个群体,单一校准无法保证评分含义一致。

Does calibration mean what they say it means; or, the reference class problem rises again

  • 个体跨多群体时,仅在某组校准不足以保证评分意义统一
  • 完美预测器才能满足所有群体的校准,现实中几乎不可能
  • 该问题影响校准及其他群体统计标准,揭示方法论盲点

关于公平性的统计标准讨论中,常以风险评分的‘含义’来强调组内校准的规范意义。在‘相同含义’图景中,不同群体的校准分数平均而言‘含义相同’,从而防止基于群体身份的差别对待。但本文认为,校准并不能保证这一点。由于具体个人属于多个群体,只有在每个人所属的所有群体中都满足校准,才能确保评分解释的一致性。然而,唯有完美预测器才可能做到这一点。因此,‘相同含义’图景犯了参考类谬误:从个体属于某一群体且该群体校准,推断其分数具有特定含义,这种推论既无依据,也极可能是错误的。随后指出,参考类问题不仅困扰校准,也影响其他声称与公平性紧密相关的群体统计标准。反思这一疏漏,揭示出算法公平性主流方法依赖简化案例的深层问题。

原文摘要 · Abstract (English)

Discussions of statistical criteria for fairness commonly convey the normative significance of calibration within groups by invoking what risk scores "mean." On the Same Meaning picture, group-calibrated scores "mean the same thing" (on average) across individuals from different groups and accordingly, guard against disparate treatment of individuals based on group membership. My contention is that calibration guarantees no such thing. Since concrete actual people belong to many groups, calibration cannot ensure the kind of consistent score interpretation that the Same Meaning picture implies matters for fairness, unless calibration is met within every group to which an individual belongs. Alas only perfect predictors may meet this bar. The Same Meaning picture thus commits a reference class fallacy by inferring from calibration within some group to the "meaning" or evidential value of an individual's score, because they are a member of that group. The reference class answer it presumes does not only lack justification; it is very likely wrong. I then show that the reference class problem besets not just calibration but other group statistical criteria that claim a close connection to fairness. Reflecting on the origins of this oversight opens a wider lens onto the predominant methodology in algorithmic fairness based on stylized cases.

公平性校准参考类谬误

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。