arXiv:2410.04120cs.LGcs.CY2024-10ICLR被引 8

揭示公平表示学习在敏感任务中的局限性,挑战现有评估方式。

Rethinking Fair Representation Learning for Performance-Sensitive Tasks

  • 用因果推理分析数据偏见来源,发现方法隐含假设
  • 实验显示分布偏移下公平学习性能显著下降
  • 适合关注医疗等高风险场景中公平性有效性的研究者

我们研究了主流的公平表示学习方法在偏见缓解中的应用。通过因果推理定义并形式化数据集偏见的不同来源,揭示了这些方法内在的重要隐含假设。证明当评估数据与训练数据来自同一分布时,公平表示学习存在根本性局限,并在多种医学模态上进行实验,考察其在分布偏移下的表现。结果解释了现有文献中的看似矛盾现象,揭示了常被忽视的因果与统计因素如何影响公平表示学习的有效性。我们对当前评估实践提出质疑,并质疑该类方法在性能敏感场景中的适用性。我们认为,未来应加强对数据偏见的细粒度分析。

原文摘要 · Abstract (English)

We investigate the prominent class of fair representation learning methods for bias mitigation. Using causal reasoning to define and formalise different sources of dataset bias, we reveal important implicit assumptions inherent to these methods. We prove fundamental limitations on fair representation learning when evaluation data is drawn from the same distribution as training data and run experiments across a range of medical modalities to examine the performance of fair representation learning under distribution shifts. Our results explain apparent contradictions in the existing literature and reveal how rarely considered causal and statistical aspects of the underlying data affect the validity of fair representation learning. We raise doubts about current evaluation practices and the applicability of fair representation learning methods in performance-sensitive settings. We argue that fine-grained analysis of dataset biases should play a key role in the field moving forward.

公平学习因果推理医疗AI分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。