arXiv:2511.09672cs.LG2025-11被引 2

用生成网络提升私有数据合成效率,百列数据仍高效运行。

GEM+: Scalable State-of-the-Art Private Synthetic Data with Generator Networks

  • 结合自适应测量与生成网络,动态优化隐私保护下的数据质量。
  • 在超百列数据上表现超越现有方法,内存与计算开销显著降低。
  • 适合处理高维真实数据,对工业级隐私数据合成极具价值。

当前最先进的差分隐私合成表格数据依赖于自适应的‘选择-测量-生成’框架(如AIM),通过迭代测量低阶噪声边际并拟合图模型生成数据,实现隐私约束下数据质量的系统优化。然而,图模型在高维数据中效率低下,需大量内存且结构变化时必须从头训练,带来巨大计算开销。近期方法如GEM采用生成神经网络提升可扩展性,但实证评估多集中于小规模数据集,限制了实际应用。本文提出GEM+,将AIM的自适应测量框架与GEM的可扩展生成网络相结合。实验表明,GEM+在数据效用与可扩展性上均优于AIM,能高效处理超过一百列的数据集,而传统AIM因内存与计算瓶颈无法胜任。

原文摘要 · Abstract (English)

State-of-the-art differentially private synthetic tabular data has been defined by adaptive 'select-measure-generate' frameworks, exemplified by methods like AIM. These approaches iteratively measure low-order noisy marginals and fit graphical models to produce synthetic data, enabling systematic optimisation of data quality under privacy constraints. Graphical models, however, are inefficient for high-dimensional data because they require substantial memory and must be retrained from scratch whenever the graph structure changes, leading to significant computational overhead. Recent methods, like GEM, overcome these limitations by using generator neural networks for improved scalability. However, empirical comparisons have mostly focused on small datasets, limiting real-world applicability. In this work, we introduce GEM+, which integrates AIM's adaptive measurement framework with GEM's scalable generator network. Our experiments show that GEM+ outperforms AIM in both utility and scalability, delivering state-of-the-art results while efficiently handling datasets with over a hundred columns, where AIM fails due to memory and computational overheads.

隐私合成生成模型差分隐私高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。