用因果推断重新定义成员推理攻击评估,无需重训即可准确测隐私风险。
Causal Evaluation of Membership Inference Attacks
- 将成员推理攻击视为因果效应,量化数据加入训练集的影响。
- 发现单次训练法存在点间干扰,事后评估受数据分布偏移影响。
- 提出三种场景下可靠估计算法,适用于大模型隐私评估。
成员推理攻击(MIA)旨在区分训练数据(成员)与未见数据(非成员),广泛用于量化模型记忆现象并评估隐私风险。传统MIA评估需多次重训练,对大模型成本高昂。单次训练(随机包含数据)和事后评估(零次训练)方法常被采用,但其统计有效性不明确。本文将MIA评估建模为因果推断问题,将‘记忆’定义为数据点纳入训练集的因果效应。该新范式揭示并形式化了现有方法的关键偏差:单次训练法存在联合包含点间的干扰,而事后评估还受成员与非成员评估数据分布偏移的混杂影响。我们推导了标准MIA指标的因果对应物,并提出了多轮、单轮与零轮场景下的实用估计器,具备非渐近一致性保证。在预训练与微调大语言模型等多个场景中验证了该方法的有效性,证明其可在无需重训且存在分布偏移时实现可靠的MIA性能测量。整体框架为现代AI系统的隐私评估提供了严谨基础。
原文摘要 · Abstract (English)
Membership Inference Attacks (MIAs) aim to distinguish training points (members) from unseen data (non-members), and are widely used to quantify memorization and assess privacy risks. Standard MIA evaluation requires repeated retraining, which is computationally costly for large models. One-run (single training with randomized data inclusion) and zero-run (post hoc evaluation) methods are often used instead, but their statistical validity remains unclear. We address this gap by framing MIA evaluation as a causal inference problem, defining \emph{memorization as the causal effect of including a data point in the training set}. This novel formulation reveals and formalizes key sources of bias in existing protocols: one-run methods suffer from interference between jointly included points, while zero-run evaluations are additionally confounded by distribution shift between member and non-member evaluation data. We derive causal analogues of standard MIA metrics and propose practical estimators for multi-run, one-run, and zero-run regimes with non-asymptotic consistency guarantees. We validate our approach in several settings, including pretrained and fine-tuned LLMs, showing that it enables reliable measurement of MIA performance without retraining and under distribution shift. Overall, our framework provides a principled foundation for privacy evaluation in modern AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。