距离度量骗人,合成数据看似私密实则易被攻击
The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
- 用记录距离衡量隐私,方法简单但本质错误
- 多模型测试显示,通过距离检测的合成数据仍易遭成员推理攻击
- 适合关注数据隐私合规与真实安全评估的研究者
合成数据被广泛用于共享数据而不泄露敏感信息。尽管成员推理攻击(MIAs)被视为评估合成数据隐私性的金标准,但从业者和研究者常依赖更简单的代理指标,如最近邻距离(DCR)。该指标通过比较训练数据与生成合成数据的相似性,以及与独立真实数据集的相似性,构建二元隐私检验:若合成数据不比真实数据集更接近训练数据,则判定为私密。本文表明,尽管计算成本低,但DCR等基于距离的指标无法有效识别隐私泄露。在多个数据集及经典模型(如Baynet、CTGAN)和近期扩散模型上,经代理指标判定为私密的合成数据仍极易遭受MIAs。我们进一步发现,该二元检验和连续度量均无法反映真实的成员推理风险,且结果在不同超参数设置和样本选择方式下一致失效。最后,我们指出此类指标存在设计缺陷,并展示了一个实际中被忽略的泄露案例。本文呼吁从业者放弃代理指标,转而以MIAs作为评估合成数据隐私的严谨标准,尤其在声称数据‘法律匿名’时。
原文摘要 · Abstract (English)
Synthetic data has become an increasingly popular way to share data without revealing sensitive information. Though Membership Inference Attacks (MIAs) are widely considered the gold standard for empirically assessing the privacy of a synthetic dataset, practitioners and researchers often rely on simpler proxy metrics such as Distance to Closest Record (DCR). These metrics estimate privacy by measuring the similarity between the training data and generated synthetic data. This similarity is also compared against that between the training data and a disjoint holdout set of real records to construct a binary privacy test. If the synthetic data is not more similar to the training data than the holdout set is, it passes the test and is considered private. In this work we show that, while computationally inexpensive, DCR and other distance-based metrics fail to identify privacy leakage. Across multiple datasets and both classical models such as Baynet and CTGAN and more recent diffusion models, we show that datasets deemed private by proxy metrics are highly vulnerable to MIAs. We similarly find both the binary privacy test and the continuous measure based on these metrics to be uninformative of actual membership inference risk. We further show that these failures are consistent across different metric hyperparameter settings and record selection methods. Finally, we argue DCR and other distance-based metrics to be flawed by design and show a example of a simple leakage they miss in practice. With this work, we hope to motivate practitioners to move away from proxy metrics to MIAs as the rigorous, comprehensive standard of evaluating privacy of synthetic data, in particular to make claims of datasets being legally anonymous.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。