arXiv:2504.16864stat.MEcs.LG2025-04被引 1

传统分解方法可能错误归因群体差异,新研究揭示其漏洞并提出修正方向。

Common Functional Decompositions Can Mis-attribute Differences in Outcomes Between Populations

  • 提出基于非线性函数分解的改进框架,避免误归因
  • 发现现有方法在相同条件下仍会错误归因差异
  • 适用于关注群体公平性与因果解释的研究者

在科学与社会科学中,我们常需解释两个群体间结果差异的原因。例如,若某就业项目在不同城市效果不同,是参与者特征(协变量)差异所致,还是当地劳动力市场(给定协变量下的结果)不同?经典Kitagawa-Oaxaca-Blinder(KOB)分解假设协变量与结果间为线性关系,但真实关系可能显著非线性。现代机器学习提供了多种单群体中结果与协变量间非线性关系的分解方法。看似自然地将这些方法扩展至KOB框架,但研究发现:成功的扩展必须确保当协变量或给定协变量下的结果在两群体中相同时,不将其差异归因于这些部分。然而,我们证明即使是简单例子中,两种常用分解——函数方差分析(functional ANOVA)与累积局部效应(Accumulated Local Effects)——仍可能将相同的结果差异错误归因于给定协变量的结果。我们给出了函数方差分析误归因的刻画,并提出任意离散分解避免误归因的必要条件:若分解独立于输入分布,则不会误归因。进一步推测,任何合理且依赖协变量分布的加性分解都可能产生误归因。

原文摘要 · Abstract (English)

In science and social science, we often wish to explain why an outcome is different in two populations. For instance, if a jobs program benefits members of one city more than another, is that due to differences in program participants (particular covariates) or the local labor markets (outcomes given covariates)? The Kitagawa-Oaxaca-Blinder (KOB) decomposition is a standard tool in econometrics that explains the difference in the mean outcome across two populations. However, the KOB decomposition assumes a linear relationship between covariates and outcomes, while the true relationship may be meaningfully nonlinear. Modern machine learning boasts a variety of nonlinear functional decompositions for the relationship between outcomes and covariates in one population. It seems natural to extend the KOB decomposition using these functional decompositions. We observe that a successful extension should not attribute the differences to covariates -- or, respectively, to outcomes given covariates -- if those are the same in the two populations. Unfortunately, we demonstrate that, even in simple examples, two common decompositions -- functional ANOVA and Accumulated Local Effects -- can attribute differences to outcomes given covariates, even when they are identical in two populations. We provide a characterization of when functional ANOVA misattributes, as well as a general property that any discrete decomposition must satisfy to avoid misattribution. We show that if the decomposition is independent of its input distribution, it does not misattribute. We further conjecture that misattribution arises in any reasonable additive decomposition that depends on the distribution of the covariates.

因果推断群体差异分解方法机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。