arXiv:2411.03351cs.CRcs.AI2024-11综述被引 11

保护隐私的表格数据生成方法综述,解决敏感信息泄露问题

Tabular Data Synthesis with Differential Privacy: A Survey

  • 按统计与深度学习模型分类,覆盖集中式与分布式场景
  • 评估生成数据在隐私性、实用性与计算开销上的平衡表现
  • 适合关注数据隐私与合成技术的研究者和从业者

数据共享是协同创新的前提,使组织能够利用多样化数据集获取深层洞察。在金融科技与智能制造等实际应用中,以表格形式存在的交易数据被广泛生成与分析。然而,这些数据通常包含敏感的个人或商业信息,引发隐私担忧与合规风险。数据合成通过生成保留真实数据统计特征的虚拟数据集,消除与具体个体的直接关联来缓解该问题。但攻击者仍可能借助背景知识推断敏感信息。差分隐私提供了可证明且量化的隐私保障,因此,差分隐私下的表格数据合成成为一种有前景的隐私友好型数据共享方案。本文系统综述现有差分隐私表格数据合成方法,揭示各类生成模型在差分隐私约束下特有的挑战。按生成模型分为统计方法与基于深度学习的方法,分别讨论其在集中式与分布式环境中的实现。评估并比较各方法在效用、隐私保护与计算复杂度方面的优劣。此外,阐述多种合成数据质量评估方法,指出当前研究空白与未来方向。

原文摘要 · Abstract (English)

Data sharing is a prerequisite for collaborative innovation, enabling organizations to leverage diverse datasets for deeper insights. In real-world applications like FinTech and Smart Manufacturing, transactional data, often in tabular form, are generated and analyzed for insight generation. However, such datasets typically contain sensitive personal/business information, raising privacy concerns and regulatory risks. Data synthesis tackles this by generating artificial datasets that preserve the statistical characteristics of real data, removing direct links to individuals. However, attackers can still infer sensitive information using background knowledge. Differential privacy offers a solution by providing provable and quantifiable privacy protection. Consequently, differentially private data synthesis has emerged as a promising approach to privacy-aware data sharing. This paper provides a comprehensive overview of existing differentially private tabular data synthesis methods, highlighting the unique challenges of each generation model for generating tabular data under differential privacy constraints. We classify the methods into statistical and deep learning-based approaches based on their generation models, discussing them in both centralized and distributed environments. We evaluate and compare those methods within each category, highlighting their strengths and weaknesses in terms of utility, privacy, and computational complexity. Additionally, we present and discuss various evaluation methods for assessing the quality of the synthesized data, identify research gaps in the field and directions for future research.

隐私保护数据合成差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。