arXiv:2607.16620cs.LGcs.AI2026-07中稿 · ed

提出隐私成本公平性度量,揭示保护代价与收益的不均衡问题。

Privacy Cost as Equity Input: A Group Fairness Criterion for Differentially Private Machine Learning

  • 以隐私暴露程度为权重,衡量各群体应得收益是否匹配代价。
  • 在COMPAS数据集上发现受保护群体面临双重不利:隐私暴露更大且预测更差。
  • 无需训练影子模型,可直接用准确率计算,适合事后审计。

差分隐私(DP)被广泛用于降低机器学习系统的成员推断风险。已有研究指出,DP-SGD可能加剧不同人口群体间的准确率差异,但这种视角仅关注结果公平性。本文认为,隐私成本——即各群体承担的信息泄露——本身就是一种伤害,并引入补偿性公平框架:承担更大隐私暴露的群体应获得相应更多收益。基于此原则,提出隐私成本公平比(PCER),定义为某群体正向预测率与其群体过拟合差距的比值。根据标准成员推断界限,该过拟合差距可上界估计群体遭受推断攻击的脆弱性,使PCER成为保守的相对收益度量。PCER仅需各群体的训练与测试准确率(无需影子模型),具备实际可操作性。我们在六个基准属性组合(涵盖表格与NLP领域)上,对DP-SGD在多种隐私预算下的表现进行评估,并验证过拟合差距作为代理指标的有效性。结果显示,传统结果导向度量无法捕捉的模式:在COMPAS数据集中,受保护群体同时承受更高隐私暴露与更差预测表现,而人口均等性差距完全掩盖了这一现象。敏感性分析表明,当隐私保证极强时,两群体过拟合水平趋同于数值下限,导致基于暴露的审计失去意义。研究强调,对隐私保护系统进行公平性审计,必须考虑谁承担了保护成本,而不仅是谁获得了收益。

原文摘要 · Abstract (English)

Differential privacy (DP) is increasingly deployed to limit membership inference risk in machine-learning systems. Prior work has shown that DP-SGD can widen accuracy disparities across demographic groups, but this framing treats fairness as a purely outcome-side concern. We argue that privacy cost, the information leakage borne by each group, is itself a form of harm, and adopt a compensatory-fairness framework in which a group that involuntarily bears greater privacy exposure is owed proportionally greater benefit from the system. From this principle we derive the \emph{Privacy-Cost Equity Ratio} (PCER), a group fairness metric defined as a group's positive prediction rate normalized by its per-group overfitting gap. By a standard membership inference bound, this overfitting gap upper-bounds each group's vulnerability to inference attacks, making PCER a conservative measure of benefit relative to exposure. PCER needs only per-group train and test accuracy (no shadow models), making it a practical post-hoc audit tool. We evaluate PCER alongside standard fairness metrics across six benchmark--attribute combinations spanning tabular and NLP domains, under DP-SGD at a range of privacy budgets, and validate the overfitting-gap proxy against a direct threshold membership-inference attack. The results reveal patterns that outcome-based metrics miss. On COMPAS, PCER uncovers a persistent double disadvantage: the protected group bears both greater privacy exposure and worse predictive outcomes, something demographic parity gap masks entirely. Sensitivity analysis shows very strong privacy guarantees collapse both groups' overfitting to a numerical floor, rendering exposure-based audits uninformative in that regime. Together, these findings show that fairness audits of privacy-preserving systems must account for who bears the cost of protection, not only who benefits from its outcomes.

差分隐私公平性隐私成本审计工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。