针对因果推断设计私密合成数据,确保处理效应估计准确
Workload-Preserving Differentially Private Synthetic Data for Causal Inference via Maximum-Entropy Calibration

- 基于双重稳健估计的正交矩构建因果工作负载
- 在严格隐私预算下,平均处理效应估计误差显著降低
- 适合需要可信因果推断的医疗、政策研究场景
基于工作负载的差分隐私(DP)合成数据方法通过测量聚合查询并后处理噪声结果生成合成记录。通用工作负载可实现强分布保真度,但因果估量如平均处理效应(ATE)依赖于处理组平衡和结果矩,而通用边际不一定保留这些特性。本文提出因果工作负载:围绕双重稳健因果估计器所用正交矩设计的DP查询集。释放的工作负载可直接用于稳定矩映射估计器,或通过最大熵校准重构为可复用的合成数据;理论分析将ATE误差分解为抽样、隐私、工作负载近似、蒙特卡洛及校准五项。我们还引入自适应工作负载选择器Causal-AIM,以及噪声感知多重插补(NA+MI)方法以获取DP合成数据的置信区间。由于工作负载只需释放一次,同一份DP合成表可支持ATE、ATT及子群分析,无需额外隐私开销。实验表明,因果工作负载在严格隐私预算下对校准不确定性最为有效,而通用工作负载在隐私放宽时仍更优于点估计均方误差。核心启示是:分布保真有助于点估计精度,但有效因果推断需保留因果矩并传播差分隐私噪声,而非将合成行视作真实数据。
原文摘要 · Abstract (English)
Workload-based differentially private (DP) synthetic data methods privately measure aggregate queries and post-process the noisy answers into synthetic records. Generic workloads can achieve strong distributional fidelity, but causal estimands such as the average treatment effect (ATE) depend on treatment-arm balance and outcome moments that generic marginals need not preserve. We propose causal workloads: DP query sets designed around the orthogonal moments used by doubly robust causal estimators. The released workload can be used directly by stable moment-map estimators or reconstructed by maximum-entropy calibration into reusable synthetic data; our theory decomposes ATE error into sampling, privacy, workload-approximation, Monte Carlo, and calibration terms. We also introduce Causal-AIM, an adaptive workload selector, and a noise-aware multiple-imputation (NA+MI) procedure for confidence intervals from DP synthetic data. Because the workload is released once, the same DP synthetic table can support ATE, ATT, and subgroup analyses without additional privacy spending. Empirically, causal workloads are most useful at strict privacy budgets and for calibrated uncertainty, while generic workloads often retain an advantage for point RMSE as privacy relaxes. The broader lesson is a tradeoff: distributional fidelity can help point accuracy, but valid causal inference requires preserving causal moments and propagating DP noise rather than treating synthetic rows as real.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。