arXiv:2508.06647cs.LG2025-08被引 2

用自回归方法生成高保真表格数据,兼顾隐私与实用

Privacy-Preserving Tabular Synthetic Data Generation Using TabularARGN

  • 基于离散化自回归结构生成表格数据,效率高
  • 在统计相似性、模型性能和抗检测上表现优异
  • 通过会员推断攻击测试,验证隐私保护能力

合成数据生成已成为安全共享和分析敏感数据集的关键手段。传统匿名化技术往往难以充分保障隐私。我们提出专为生成高质量表格数据设计的表格式自回归生成网络(TabularARGN),采用基于离散化的自回归方法,在保持高数据保真度的同时具备良好计算效率。我们在多项指标上对比现有方法,结果表明其在统计相似性、机器学习可用性及抗检测能力方面具有竞争力。进一步通过系统性的成员推断攻击评估隐私性,验证了该方法在隐私与效用间取得良好平衡。

原文摘要 · Abstract (English)

Synthetic data generation has become essential for securely sharing and analyzing sensitive data sets. Traditional anonymization techniques, however, often fail to adequately preserve privacy. We introduce the Tabular Auto-Regressive Generative Network (TabularARGN), a neural network architecture specifically designed for generating high-quality synthetic tabular data. Using a discretization-based auto-regressive approach, TabularARGN achieves high data fidelity while remaining computationally efficient. We evaluate TabularARGN against existing synthetic data generation methods, showing competitive results in statistical similarity, machine learning utility, and detection robustness. We further perform an in-depth privacy evaluation using systematic membership-inference attacks, highlighting the robustness and effective privacy-utility balance of our approach.

合成数据隐私保护表格生成自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。