跨处理水平融合信息,提升数据稀疏时的因果矩阵补全效果
Causal Matrix Completion under Multiple Treatments via Mixed Synthetic Nearest Neighbors
- 通过混合不同处理水平的近邻信息,扩展可用样本量
- 在小样本处理组下仍保持误差有界和渐近正态性
- 适合处理多类或复杂干预场景下的缺失数据问题
合成最近邻(SNN)通过利用完全观测锚定子矩阵的局部低秩结构,为缺失非随机(MNAR)场景下的因果矩阵补全提供了理论完备的解决方案。然而其有效性高度依赖各处理水平内充足的数据量,这在多处理或复杂处理设置中常难以满足。本文提出混合合成最近邻(MSNN),一种新的逐元素因果识别估计器,能够整合不同处理水平的信息。我们证明,MSNN保持了SNN的有限样本误差界和渐近正态性保证,同时扩大了可用于估计的有效样本规模。在合成与真实数据集上的实证结果表明,该方法在处理水平数据稀缺时尤为有效。
原文摘要 · Abstract (English)
Synthetic Nearest Neighbors (SNN) provides a principled solution to causal matrix completion under missing-not-at-random (MNAR) by exploiting local low-rank structure through fully observed anchor submatrices. However, its effectiveness critically relies on sufficient data availability within each treatment level, a condition that often fails in settings with multiple or complex treatments. In this work, we propose Mixed Synthetic Nearest Neighbors (MSNN), a new entry-wise causal identification estimator that integrates information across treatment levels. We show that MSNN retains the finite-sample error bounds and asymptotic normality guarantees of SNN, while enlarging the effective sample size available for estimation. Empirical results on synthetic and real-world datasets illustrate the efficacy of the proposed approach, especially under data-scarce treatment levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。