arXiv:2506.06108cs.LGcs.CR2025-06中稿 · KDD被引 14

综述表格型合成数据生成、攻击与防御方法

Synthetic Tabular Data: Methods, Attacks and Defenses

  • 基于概率图模型与深度学习的生成范式
  • 揭示合成数据泄露原始隐私信息的风险
  • 适合关注数据隐私与生成模型的研究者

合成数据常被视为替代敏感固定规模数据集的方案,可提供无限量且无隐私顾虑的匹配数据。过去十年间,随着机器学习与数据分析技术的发展,合成数据生成取得了显著进展。本文综述了表格型合成数据生成的关键进展与核心概念,涵盖基于概率图模型与深度学习的生成范式。在介绍背景与动机后,深入剖析技术方法。同时,通过研究旨在恢复原始敏感数据信息的攻击,揭示合成数据的局限性。最后,讨论该领域的扩展方向与开放问题。

原文摘要 · Abstract (English)

Synthetic data is often positioned as a solution to replace sensitive fixed-size datasets with a source of unlimited matching data, freed from privacy concerns. There has been much progress in synthetic data generation over the last decade, leveraging corresponding advances in machine learning and data analytics. In this survey, we cover the key developments and the main concepts in tabular synthetic data generation, including paradigms based on probabilistic graphical models and on deep learning. We provide background and motivation, before giving a technical deep-dive into the methodologies. We also address the limitations of synthetic data, by studying attacks that seek to retrieve information about the original sensitive data. Finally, we present extensions and open problems in this area.

合成数据隐私保护表格生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。