用切尔诺夫信息分析数据隐私与公平的权衡关系
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
- 提出切尔诺夫差异衡量数据公平性,引入噪声变体统一分析公平与隐私
- 在高斯示例中发现三类不同行为模式,揭示依赖数据分布的本质权衡
- 开发首个神经网络估计器CINE,实现真实数据上公平-隐私关系的量化分析
公平性和隐私性是可信机器学习的两大支柱。尽管相关研究广泛,但二者关系仍被严重忽视。本文利用信息论中的切尔诺夫信息,刻画输入数据分布引发的公平性、隐私性与准确性的基本权衡。我们提出切尔诺夫差异(Chernoff Difference)作为数据公平性的度量,并引入其噪声版本——噪声切尔诺夫差异(Noisy Chernoff Difference),以同时分析公平与隐私。通过简单的高斯例子,我们发现噪声切尔诺夫差异的行为在不同数据分布下呈现三种定性不同的模式。为将分析扩展至真实场景,我们提出了切尔诺夫信息神经估计器(CINE),这是首个用于未知分布的切尔诺夫信息神经网络估计方法。我们将CINE应用于真实数据集,系统分析了噪声切尔诺夫差异。本工作填补了该领域关键空白,提供了基于数据的、原则性的公平-隐私交互机制分析。
原文摘要 · Abstract (English)
Fairness and privacy are two vital pillars of trustworthy machine learning. Despite extensive research on these individual topics, their relationship has received significantly less attention. In this paper, we utilize an information-theoretic measure Chernoff Information to characterize the fundamental trade-off between fairness, privacy, and accuracy, as induced by the input data distribution. We first propose Chernoff Difference, a notion of data fairness, along with its noisy variant, Noisy Chernoff Difference, which allows us to analyze both fairness and privacy simultaneously. Through simple Gaussian examples, we show that Noisy Chernoff Difference exhibits three qualitatively distinct behaviors depending on the underlying data distribution. To extend this analysis beyond synthetic settings, we develop the Chernoff Information Neural Estimator (CINE), the first neural network-based estimator of Chernoff Information for unknown distributions. We apply CINE to analyze the Noisy Chernoff Difference on real-world datasets. Together, this work fills a critical gap in the literature by providing a principled, data-dependent characterization of the fairness-privacy interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。