提出可保护隐私的比率统计方法,小样本下仍有效。
Differentially private ratio statistics
- 设计简单算法实现隐私与准确性的良好平衡
- 证明相对风险估计器的一致性并构造有效置信区间
- 适合需隐私保护的因果推断与公平性分析场景
比率统计(如相对风险、优势比)在机器学习中的假设检验、模型评估和决策中起核心作用,尤其在因果推断与公平性分析中。然而,尽管数据集普遍存在隐私顾虑且差分隐私日益普及,现有文献对差分隐私下的比率统计关注甚少,仅林等[1]近期有初步研究。本文旨在填补这一空白,提供可指导实践的比率估计方法,确保结果受差分隐私保护。我们表明,即使简单算法在小样本下也能同时具备良好的隐私性、样本准确性和低偏差。此外,我们分析了一种差分隐私下的相对风险估计器,证明其一致性,并提出构造有效置信区间的办法。该方法弥合了差分隐私文献中的关键缺口,为私有机器学习流程中的比率估计提供了实用解决方案。
原文摘要 · Abstract (English)
Ratio statistics--such as relative risk and odds ratios--play a central role in hypothesis testing, model evaluation, and decision-making across many areas of machine learning, including causal inference and fairness analysis. However, despite privacy concerns surrounding many datasets and despite increasing adoption of differential privacy, differentially private ratio statistics have largely been neglected by the literature and have only recently received an initial treatment by Lin et al. [1]. This paper attempts to fill this lacuna, giving results that can guide practice in evaluating ratios when the results must be protected by differential privacy. In particular, we show that even a simple algorithm can provide excellent properties concerning privacy, sample accuracy, and bias, not just asymptotically but also at quite small sample sizes. Additionally, we analyze a differentially private estimator for relative risk, prove its consistency, and develop a method for constructing valid confidence intervals. Our approach bridges a gap in the differential privacy literature and provides a practical solution for ratio estimation in private machine learning pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。