即使释放无限多合成数据,隐私放大仍有效,突破了以往理论限制。
Privacy Amplification Persists under Unlimited Synthetic Data Release
- 在参数有界假设下,合成数据释放仍可实现隐私放大。
- 突破了此前仅适用于高维模型与少量数据的理论局限。
- 为复杂数据发布机制的隐私保障提供新思路,适合关注隐私计算的研究者。
我们研究了合成数据释放带来的隐私放大现象,即仅发布合成数据而非私有生成模型时,差分隐私保证得以增强。Pierquin 等人(2025)首次为线性生成器建立了形式化放大保证,但这些结果仅在模型维度远大于释放的合成记录数的渐近情形下成立,限制了实际应用价值。本文证明了一个令人惊讶的结果:在参数有界假设下,即使释放任意数量的合成记录,隐私放大依然存在,从而改进了 Pierquin 等人(2025)的边界。我们的分析揭示了结构特性,可能指导更复杂释放机制下更紧致隐私保证的开发。
原文摘要 · Abstract (English)
We study privacy amplification by synthetic data release, a phenomenon in which differential privacy guarantees are improved by releasing only synthetic data rather than the private generative model itself. Recent work by Pierquin et al. (2025) established the first formal amplification guarantees for a linear generator, but they apply only in asymptotic regimes where the model dimension far exceeds the number of released synthetic records, limiting their practical relevance. In this work, we show a surprising result: under a bounded-parameter assumption, privacy amplification persists even when releasing an unbounded number of synthetic records, thereby improving upon the bounds of Pierquin et al. (2025). Our analysis provides structural insights that may guide the development of tighter privacy guarantees for more complex release mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。