arXiv:2501.03941cs.LGcs.AI2025-01被引 13

评估合成数据隐私的多种方法,助力安全数据共享。

Synthetic Data Privacy Metrics

  • 对比主流隐私度量方法,分析其优劣。
  • 提出改进生成模型以增强数据隐私的最佳实践。
  • 适合关注数据安全与隐私保护的研究者参考。

生成式AI的进步使得合成数据集能够达到真实数据的精度,可用于训练人工智能模型、提供统计洞察,并在敏感数据协作中提供强隐私保障。有效衡量合成数据的实证隐私是这一过程中的关键步骤。然而,尽管每天都有大量新隐私度量方法发表,目前仍缺乏统一标准。本文综述了包括对抗攻击模拟在内的主流隐私度量方法的优缺点,并回顾了当前提升生成模型所产数据隐私性的最佳实践(如差分隐私)。

原文摘要 · Abstract (English)

Recent advancements in generative AI have made it possible to create synthetic datasets that can be as accurate as real-world data for training AI models, powering statistical insights, and fostering collaboration with sensitive datasets while offering strong privacy guarantees. Effectively measuring the empirical privacy of synthetic data is an important step in the process. However, while there is a multitude of new privacy metrics being published every day, there currently is no standardization. In this paper, we review the pros and cons of popular metrics that include simulations of adversarial attacks. We also review current best practices for amending generative models to enhance the privacy of the data they create (e.g. differential privacy).

隐私保护合成数据差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。