揭示表格扩散模型的隐私泄露风险及关键影响因素。
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics

- 通过黑盒与白盒攻击量化训练设置、生成策略和攻击者知识的影响。
- 发现攻击者无需完美信息或大量算力即可成功发起攻击。
- 指出常用隐私度量方法(如最近邻距离)存在明显缺陷。
表格数据在众多领域中至关重要,尤其涉及高隐私风险场景。为降低隐私泄露与敏感数据暴露风险,生成高质量合成表格数据成为重要手段。当前表格扩散模型(TDMs)在数据合成方面表现领先,但其潜在隐私风险亟需深入理解与评估。本研究利用先进的成员推断攻击技术,在黑盒与白盒设定下,量化分析了训练配置、合成策略及攻击者知识对隐私泄露的影响。结果表明,攻击者无需完全掌握训练细节、相同数据分布或大规模计算资源即可成功实施攻击。此外,研究揭示了诸如‘最近邻距离’等启发式隐私度量方法存在的严重问题,提示现有评估方式可能产生误导性结论。
原文摘要 · Abstract (English)
Tabular data plays an important role in many fields and industries, including those with elevated privacy considerations and risks. As such, there is a rising interest in generating high-quality synthetic proxies for real tabular data as a means of reducing privacy risk and proprietary data exposure. With tabular diffusion models (TDMs) demonstrating leading performance in synthesizing such data, understanding and measuring the privacy risks associated with these models is imperative. Leveraging state-of-the-art membership inference attacks for TDMs in both black- and white-box settings, this work quantifies the impact of training setup, synthesis choices, and attacker knowledge on privacy leakage. Moreover, the results demonstrate that adversaries need not have perfect knowledge of the training setup, identical data distributions, or massive compute resources to construct successful attacks. Finally, the pitfalls associated with applying heuristic privacy metrics, such as distance-to-closest record, are revealed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。