arXiv:2604.27741cs.LG2026-04

找出两组人群在特征空间中的差异点及其原因。

Differential Subgroup Discovery: Characterizing Where Two Populations Differ, and Why

论文配图:Differential Subgroup Discovery: Characterizing Where Two Populations Differ, and Why
图 1 · 摘自论文原文
  • 基于优化目标寻找具有显著结果差异的子群体。
  • 在医学、模型诊断等场景中准确识别差异区域。
  • 可解释性强,适合临床分析与因果推断研究者使用。

我们研究了如何理解两个群体在特征空间中的差异,提出“差异子群”概念:即在特征上相似但目标结果差异显著的个体集合。这些子群揭示了群体间差距最明显的区域,并帮助识别导致差异的协变量组合,适用于临床分析、模型诊断和因果效应研究。本文提出一个通用优化目标,并建立其在特定条件下具备因果解释性的理论基础。进而设计了梯度驱动的DiffSub方法,用于在表格数据中发现可解释的差异子群。在合成基准、医疗案例、模型误差分析及治疗效果评估中,DiffSub均能有效识别出有意义的子群体,揭示差异发生的位置与成因。

原文摘要 · Abstract (English)

We study the problem of understanding where two populations differ within a feature space, which we formalize in the concept of a differential subgroup: a subset of individuals from both populations who, despite sharing similar characteristics, exhibit exceptional differences in a target outcome. Differential subgroups reveal the regions of the feature space where population-level gaps are most pronounced and can help practitioners identify the covariate combinations that are structurally responsible for these differences, e.g.~in clinical analysis, model diagnostics, or treatment-effect studies. We introduce a general optimization objective for discovering differential subgroups and establish conditions under which the resulting subgroups admit a causal interpretation of population differences. We propose DiffSub, a gradient-based approach that discovers interpretable differential subgroups in tabular data. Across synthetic benchmarks, medical case studies, model-error analyses, and treatment-effect settings, DiffSub identifies informative subgroups that reveal where population differences arise and why.

子群发现因果分析可解释性统计差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。