提出审计方法,检测数据注入如何改变因果效应估计结果。
Data-Poisoning Audits for Causal Effect Estimation
- 构建可计算最坏情况变化的贪心扫描算法。
- 发现小规模数据注入即可显著影响因果估计结果。
- 适合关注数据安全与因果推断可靠性的研究者使用。
观测性因果分析日益整合来自不同来源、供应商和采集系统的数据,使其面临仅追加型攻击的威胁:攻击者精心选择看似合理的记录以改变报告的处理效应。本文针对增广逆概率加权估计开发了数据注入审计方法。分析师指定可行记录的有限目录、追加预算及嵌套源容量,攻击者则选择可行子集以最大化预设方向上的效应变动。在预处理和扰动模型固定的情况下,提出一种贪心扫描算法,可在每个追加预算下精确计算有限样本下的最坏情况变动。为考虑扰动重拟合的影响,进一步推导出结合每条记录直接贡献及其通过倾向得分和结果模型影响的总影响分数,并获得完全重拟合估计的保守有限预算边界。大量模拟验证了精确结果,显示总影响分数提升局部重拟合预测效果;多源及公开数据实证表明,在小追加预算下即存在显著敏感性。该框架将对抗性数据组合风险转化为移动曲线与临界预算,支持更可靠的因果报告并促进源头防护设计。
原文摘要 · Abstract (English)
Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect. We develop a data-poisoning audit for augmented inverse-probability-weighted estimation. The analyst specifies a finite catalog of feasible records, an append budget, and nested source capacities, and the adversary selects a feasible subset to maximize movement in a prespecified direction. With preprocessing and nuisance fits held fixed, we propose a greedy scan that computes the exact finite-sample worst-case movement at every append budget. To account for nuisance refitting, we go on to derive a total-influence score combining each record's direct contribution with its effect through the propensity and outcome models. We further obtain a conservative finite-budget bound for the fully refitted estimate. Extensive simulations validate the exact result and show that total influence improves local refit prediction, while multisite and public-data analyses demonstrate material sensitivity at small append budgets. By translating adversarial data-composition risk into movement curves and critical budgets, the framework supports more reliable causal reporting and the design of source-level safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。