arXiv:2505.19589cs.LGstat.ML2025-05被引 3

无需强假设,用预测扰动实现隐私保护的因果推断新方法。

Model Agnostic Differentially Private Causal Inference

  • 分离扰动与模型估计,仅对预测和聚合步骤加噪。
  • 在真实隐私预算下,三种经典估计器性能保持竞争力。
  • 适合医疗、经济等需隐私保护的因果分析场景。

从观测数据中估计因果效应在医学、经济学和社会科学中至关重要,但隐私问题尤为突出。本文提出一种通用且模型无关的差分隐私平均处理效应(ATE)估计框架,避免对数据生成过程或倾向得分与条件结果模型施加强结构假设。与以往直接对冗余成分进行隐私化的方法不同,本方法将冗余估计与隐私保护解耦:使用灵活的黑箱模型进行估计,而差分隐私通过折叠分割方案结合集成技术,仅对预测和聚合步骤加噪实现。该框架应用于三种经典估计器——G-公式、逆倾向得分加权(IPW)和增广IPW(AIPW),提供形式化的效用与隐私保证,并给出私有化置信区间。在合成与真实数据上的实验表明,该方法在合理隐私预算下仍保持良好性能。

原文摘要 · Abstract (English)

Estimating causal effects from observational data is essential in fields such as medicine, economics and social sciences, where privacy concerns are paramount. We propose a general, model-agnostic framework for differentially private estimation of average treatment effects (ATE) that avoids strong structural assumptions on the data-generating process or the models used to estimate propensity scores and conditional outcomes. In contrast to prior work, which enforces differential privacy by directly privatizing these nuisance components, our approach decouples nuisance estimation from privacy protection. This separation allows the use of flexible, state-of-the-art black-box models, while differential privacy is achieved by perturbing only predictions and aggregation steps within a fold-splitting scheme with ensemble techniques. We instantiate the framework for three classical estimators -- the G-Formula, inverse propensity weighting (IPW), and augmented IPW (AIPW) -- and provide formal utility and privacy guarantees, together with privatized confidence intervals. Empirical results on synthetic and real data show that our methods maintain competitive performance under realistic privacy budgets.

因果推断差分隐私模型无关隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。