arXiv:2504.10987cs.LGcs.CR2025-04中稿 · the Synthetic Data…被引 2

利用公共属性提升私有数据合成质量,解决隐私与精度的矛盾。

Leveraging Vertical Public-Private Split for Improved Synthetic Data Generation

  • 将横向公共数据辅助转为纵向属性划分,更贴合真实数据结构。
  • 在有限公共属性下,合成数据统计特性保持良好,但质量仍有提升空间。
  • 适合需要隐私保护且部分字段可公开的医疗、金融等场景。

差分隐私合成数据生成(DP-SDG)是实现安全表格式数据共享的关键技术,通过添加精确校准的统计噪声来保障个体隐私,但会牺牲合成数据的质量。近期研究探索了使用少量公共数据提升合成数据质量的场景,这些方法通常采用横向公共-私有划分,即利用少量公共行进行模型初始化,仅带来轻微性能增益。然而,现实数据常天然包含公共与私有属性,使纵向公共-私有划分更具实际意义。本文提出一种新框架,将横向公共数据辅助方法适配至纵向设置,并与基于条件生成的替代方案对比,揭示现有公共数据辅助方法的初始局限性,提出未来研究方向。

原文摘要 · Abstract (English)

Differentially Private Synthetic Data Generation (DP-SDG) is a key enabler of private and secure tabular-data sharing, producing artificial data that carries through the underlying statistical properties of the input data. This typically involves adding carefully calibrated statistical noise to guarantee individual privacy, at the cost of synthetic data quality. Recent literature has explored scenarios where a small amount of public data is used to help enhance the quality of synthetic data. These methods study a horizontal public-private partitioning which assumes access to a small number of public rows that can be used for model initialization, providing a small utility gain. However, realistic datasets often naturally consist of public and private attributes, making a vertical public-private partitioning relevant for practical synthetic data deployments. We propose a novel framework that adapts horizontal public-assisted methods into the vertical setting. We compare this framework against our alternative approach that uses conditional generation, highlighting initial limitations of public-data assisted methods and proposing future research directions to address these challenges.

合成数据差分隐私纵向划分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。