在保护隐私的前提下,实现观察数据中的因果推断。
Differentially Private Covariate Balancing Causal Inference
- 两阶段加权法在差分隐私约束下保持协变量平衡。
- 在给定隐私预算下,估计结果具有一致性和最优率收敛性。
- 适合医疗、金融等需保护敏感数据的因果分析场景。
差分隐私是当前主流的隐私保护数学框架,通过向原始数据添加随机化算法来提供概率保障,防止个体隐私信息泄露。然而,这种随机化会扭曲数据内在模式,给数据分析带来挑战。在隐私敏感场景中,利用观测数据进行因果推断尤为困难,因为需要确保处理组间的协变量平衡,但直接检查真实协变量又可能造成敏感信息泄露。本文提出一种差分隐私下的两阶段协变量平衡加权估计方法,可在给定隐私预算下生成具有统计保证(如一致性与率最优性)的点估计和区间估计,实现对因果效应的可靠推断。
原文摘要 · Abstract (English)
Differential privacy is the leading mathematical framework for privacy protection, providing a probabilistic guarantee that safeguards individuals' private information when publishing statistics from a dataset. This guarantee is achieved by applying a randomized algorithm to the original data, which introduces unique challenges in data analysis by distorting inherent patterns. In particular, causal inference using observational data in privacy-sensitive contexts is challenging because it requires covariate balance between treatment groups, yet checking the true covariates is prohibited to prevent leakage of sensitive information. In this article, we present a differentially private two-stage covariate balancing weighting estimator to infer causal effects from observational data. Our algorithm produces both point and interval estimators with statistical guarantees, such as consistency and rate optimality, under a given privacy budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。