提出可持续更新的私有高维表格数据生成框架
Differentially Private Synthetic High-dimensional Tabular Stream
- 基于选择-测量-拟合-迭代范式与私有计数器实现持续更新
- 在真实数据集上验证了方法在高维表格数据上的有效性
- 适合需长期追踪数据变化的隐私敏感场景
尽管差分隐私合成数据生成已在文献中广泛研究,但当底层私有数据发生变化时如何在未来更新这些数据仍知之甚少。本文提出一种算法框架,用于流式数据,能够随时间生成多个合成数据集,以追踪底层私有数据的变化。该算法满足整个输入流的差分隐私(持续差分隐私),并适用于高维表格数据。此外,我们通过在真实世界数据集上的实验展示了方法的有效性。所提出的算法基于流行的‘选择-测量-拟合-迭代’范式(常用于离线合成数据生成算法)以及用于流数据的私有计数器。
原文摘要 · Abstract (English)
While differentially private synthetic data generation has been explored extensively in the literature, how to update this data in the future if the underlying private data changes is much less understood. We propose an algorithmic framework for streaming data that generates multiple synthetic datasets over time, tracking changes in the underlying private data. Our algorithm satisfies differential privacy for the entire input stream (continual differential privacy) and can be used for high-dimensional tabular data. Furthermore, we show the utility of our method via experiments on real-world datasets. The proposed algorithm builds upon a popular select, measure, fit, and iterate paradigm (used by offline synthetic data generation algorithms) and private counters for streams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。