通过聚类筛选关键对照组,提升个体层面因果推断的准确性。
ClusterSC: Advancing Synthetic Control with Donor Selection
- 先聚类再选对照组,减少高维数据干扰
- 在真实与合成数据上均显著优于传统方法
- 适合处理大规模个体级观测数据的研究者
在观察性研究的因果推断中,合成控制(SC)已成为重要工具。传统SC多用于聚合数据,近年扩展至个体层面数据。由于个体数量增多,引发维度灾难。为此,我们提出聚类合成控制(ClusterSC),基于不同群体内部行为一致、群体间差异明显的假设,引入聚类步骤,仅选择相关对照组。理论证明了该方法带来的改进,并在合成与真实数据集上验证。结果表明,ClusterSC在各项指标上持续优于经典SC方法。
原文摘要 · Abstract (English)
In causal inference with observational studies, synthetic control (SC) has emerged as a prominent tool. SC has traditionally been applied to aggregate-level datasets, but more recent work has extended its use to individual-level data. As they contain a greater number of observed units, this shift introduces the curse of dimensionality to SC. To address this, we propose Cluster Synthetic Control (ClusterSC), based on the idea that groups of individuals may exist where behavior aligns internally but diverges between groups. ClusterSC incorporates a clustering step to select only the relevant donors for the target. We provide theoretical guarantees on the improvements induced by ClusterSC, supported by empirical demonstrations on synthetic and real-world datasets. The results indicate that ClusterSC consistently outperforms classical SC approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。