用极值理论评估生成数据隐私泄露,可定位具体样本风险。
PRIVET: Privacy Metric Based on Extreme Value Theory
- 基于最近邻距离的极值统计,为每个合成样本打隐私泄露分数。
- 在高维、小样本等极端场景下仍能准确识别记忆现象。
- 适用于基因数据等敏感领域,适合需要细粒度隐私审计的场景。
深度生成模型常训练于敏感数据,如基因序列、医疗数据或受版权保护的内容,引发对合成数据隐私保护的关切,尤其是隐私泄露问题,这与过拟合密切相关。现有方法几乎仅依赖全局指标评估模型隐私失败风险,提供的是不可解释的定量结果。缺乏样本级的严格隐私评估方法,阻碍了合成数据在实际应用中的部署。本文提出 PRIVET,一种通用的、与模态无关的样本级隐私评估算法,利用极值统计分析最近邻距离,为每个合成样本分配个体隐私泄露评分。实证表明,PRIVET 在多种数据模态中可靠检测到记忆和隐私泄露,包括高维、小样本(如基因数据)甚至欠拟合情形。在受控条件下对比现有方法,PRIVET 能同时提供定性和定量的样本级与数据集级评估。此外,分析揭示现有计算机视觉嵌入在识别近似重复样本时,难以生成感知上有意义的距离。
原文摘要 · Abstract (English)
Deep generative models are often trained on sensitive data, such as genetic sequences, health data, or more broadly, any copyrighted, licensed or protected content. This raises critical concerns around privacy-preserving synthetic data, and more specifically around privacy leakage, an issue closely tied to overfitting. Existing methods almost exclusively rely on global criteria to estimate the risk of privacy failure associated to a model, offering only quantitative non interpretable insights. The absence of rigorous evaluation methods for data privacy at the sample-level may hinder the practical deployment of synthetic data in real-world applications. Using extreme value statistics on nearest-neighbor distances, we propose PRIVET, a generic sample-based, modality-agnostic algorithm that assigns an individual privacy leak score to each synthetic sample. We empirically demonstrate that PRIVET reliably detects instances of memorization and privacy leakage across diverse data modalities, including settings with very high dimensionality, limited sample sizes such as genetic data and even under underfitting regimes. We compare our method to existing approaches under controlled settings and show its advantage in providing both dataset level and sample level assessments through qualitative and quantitative outputs. Additionally, our analysis reveals limitations in existing computer vision embeddings to yield perceptually meaningful distances when identifying near-duplicate samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。