arXiv:2604.19357cs.LG2026-04中稿 · ACM FAccT 2026被引 2

提出FairTree算法,精准检测模型在不同群体中的性能差异。

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition

论文配图:FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition
图 1 · 摘自论文原文
  • 基于偏差-方差分解,直接处理连续、分类和有序特征
  • 模拟实验显示其假阳性率低,波动检验法检出力更强
  • 适合关注模型公平性的研究者在小数据场景使用

机器学习模型评估通常依赖损失函数相关的性能指标,可能忽略关键子群体的表现变化。现有审计工具如SliceFinder和SliceLine存在概念缺陷,难以直接处理连续变量。本文提出FairTree,一种源自心理测量不变性检验的新算法。与现有方法不同,FairTree无需对连续、分类或有序特征进行离散化,可直接建模。该方法将性能差异分解为系统性偏差与方差,实现对算法表现变化的分类分析。提出两种变体:基于置换的版本(概念上接近SliceFinder)和波动检验法。通过模拟研究(含与SliceLine的直接对比)表明,两者均保持较低假阳性率,但波动检验法具有更高检出力。进一步在UCI Adult Census数据集上验证了方法的有效性。所提算法为各类应用中模型性能与公平性的统计评估提供了灵活框架,尤其适用于小样本场景。

原文摘要 · Abstract (English)

The evaluation of machine learning models typically relies mainly on performance metrics based on loss functions, which risk to overlook changes in performance in relevant subgroups. Auditing tools such as SliceFinder and SliceLine were proposed to detect such groups, but usually have conceptual disadvantages, such as the inability to directly address continuous covariates. In this paper, we introduce FairTree, a novel algorithm adapted from psychometric invariance testing. Unlike SliceFinder and related algorithms, FairTree directly handles continuous, categorical, and ordinal features without discretization. It further decomposes performance disparities into systematic bias and variance, allowing a categorization of changes in algorithm performance. We propose and evaluate two variations of the algorithm: a permutation-based approach, which is conceptually closer to SliceFinder, and a fluctuation test. Through simulation studies that include a direct comparison with SliceLine, we demonstrate that both approaches have a satisfactory rate of false-positive results, but that the fluctuation approach has relatively higher power. We further illustrate the method on the UCI Adult Census dataset. The proposed algorithms provide a flexible framework for the statistical evaluation of the performance and aspects of fairness of machine learning models in a wide range of applications even in relatively small data.

模型公平性偏差分析子群体检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。