arXiv:2412.16083cs.LGq-fin.ST2024-12被引 2

用差分隐私保护联邦学习中的表格数据生成,兼顾隐私与数据质量。

Federated Diffusion Modeling with Differential Privacy for Tabular Data Synthesis

  • 结合差分隐私、联邦学习与去噪扩散模型,实现隐私保护下的表格合成。
  • 在多个真实数据集上验证,提升隐私保障且不降低数据质量。
  • 适合金融、医疗等高监管领域安全共享数据的场景。

随着各领域对隐私保护数据分析需求的增长,亟需能严格遵守隐私标准的合成数据生成方案。我们提出 DP-FedTabDiff 框架,首次将差分隐私、联邦学习与去噪扩散概率模型融合,用于生成高质量的合成表格数据。该框架在满足隐私合规要求的同时保持数据可用性。我们在多个真实世界混合类型表格数据集上验证了其有效性,显著提升了隐私保障水平,且未牺牲数据质量。实验结果揭示了隐私预算、客户端配置与联邦优化策略之间的最优权衡关系。研究证实,DP-FedTabDiff 有望在高度监管领域推动安全数据共享与分析,为联邦学习与隐私保护数据合成的进一步发展奠定基础。

原文摘要 · Abstract (English)

The increasing demand for privacy-preserving data analytics in various domains necessitates solutions for synthetic data generation that rigorously uphold privacy standards. We introduce the DP-FedTabDiff framework, a novel integration of Differential Privacy, Federated Learning and Denoising Diffusion Probabilistic Models designed to generate high-fidelity synthetic tabular data. This framework ensures compliance with privacy regulations while maintaining data utility. We demonstrate the effectiveness of DP-FedTabDiff on multiple real-world mixed-type tabular datasets, achieving significant improvements in privacy guarantees without compromising data quality. Our empirical evaluations reveal the optimal trade-offs between privacy budgets, client configurations, and federated optimization strategies. The results affirm the potential of DP-FedTabDiff to enable secure data sharing and analytics in highly regulated domains, paving the way for further advances in federated learning and privacy-preserving data synthesis.

隐私计算联邦学习数据合成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。