数据偏差下分组评估易误导,需因果分析补救
Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
- 用因果图模型分析不同数据生成过程下的公平性
- 真实分布中各组表现相等不代表公平
- 适合关注算法公平性的研究者与实践者
对子群体进行细分评估是检验机器学习模型公平性的关键,但其未经批判地使用可能误导实践者。当数据反映真实世界不平等且代表相关人群时,各组表现一致并不能可靠衡量公平性。若因选择偏差导致数据不具代表性,则基于条件独立性检验的替代方法也可能无效,除非明确假设偏差机制。我们利用因果图模型,在不同数据生成过程中刻画公平性属性与度量稳定性。该框架建议在细分评估基础上,加入显式因果假设与分析,以控制混杂因素和分布偏移,包括条件独立性检验与加权性能估计。这些发现对普遍采用细分评估的模型评估设计与解读具有广泛影响。
原文摘要 · Abstract (English)
Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal performance across subgroups is an unreliable measure of fairness when data are representative of the relevant populations but reflective of real-world disparities. Furthermore, when data are not representative due to selection bias, both disaggregated evaluation and alternative approaches based on conditional independence testing may be invalid without explicit assumptions regarding the bias mechanism. We use causal graphical models to characterize fairness properties and metric stability across subgroups under different data generating processes. Our framework suggests complementing disaggregated evaluations with explicit causal assumptions and analysis to control for confounding and distribution shift, including conditional independence testing and weighted performance estimation. These findings have broad implications for how practitioners design and interpret model assessments given the ubiquity of disaggregated evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。