arXiv:2506.05101cs.LGcs.CR2025-06ICML被引 5

用线性回归证明:生成数据可增强隐私,但依赖模型隐藏程度。

Privacy Amplification Through Synthetic Data: Insights from Linear Regression

  • 基于线性回归分析合成数据的隐私机制。
  • 仅释放少量合成数据点时,隐私保护效果优于模型本身。
  • 适合关注隐私增强与生成模型安全的研究者。

合成数据继承生成模型的差分隐私保障。此外,当生成模型保持隐藏时,合成数据可能获得隐私放大效应。尽管经验研究已暗示此现象,但缺乏严格的理论支持。本文通过线性回归这一清晰框架展开研究:首先,证明若攻击者掌控生成模型的随机种子,单个合成数据点泄露的信息量可等同于直接发布模型本身;反之,当合成数据由随机输入生成时,有限数量的合成数据释放可实现超越模型原始保障的隐私放大。我们认为,线性回归中的发现可为未来推导更通用边界提供基础。

原文摘要 · Abstract (English)

Synthetic data inherits the differential privacy guarantees of the model used to generate it. Additionally, synthetic data may benefit from privacy amplification when the generative model is kept hidden. While empirical studies suggest this phenomenon, a rigorous theoretical understanding is still lacking. In this paper, we investigate this question through the well-understood framework of linear regression. First, we establish negative results showing that if an adversary controls the seed of the generative model, a single synthetic data point can leak as much information as releasing the model itself. Conversely, we show that when synthetic data is generated from random inputs, releasing a limited number of synthetic data points amplifies privacy beyond the model's inherent guarantees. We believe our findings in linear regression can serve as a foundation for deriving more general bounds in the future.

差分隐私合成数据线性回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。