用分离模型生成表格数据,提升隐私与计算效率。
Disjoint Generation of Synthetic Data
- 将数据切分后由独立模型分别生成,再无标识符拼接。
- 隐私性提升,重识别风险显著降低,下游任务准确率高。
- 支持混合模型生成,兼顾隐私与数据效用,适合敏感数据处理。
我们提出一种基于分离生成模型的表格合成数据新框架。该方法将数据集划分为互不重叠的子集,分别交由不同的生成模型处理,随后通过无需共用变量或标识符的拼接操作合并结果。多个案例研究验证了该框架的有效性,展示了设计选择的影响。其优势包括:(i) 实证测量的隐私性明显提升;(ii) 某些模型类型的计算可行性增强;(iii) 可融合多种生成模型进行合成。混合模型合成有效弥合隐私与效用之间的差距,在下游任务中实现高精度和高曲线下面积(AUC),同时显著降低实际重识别风险。
原文摘要 · Abstract (English)
We propose a new framework for generating tabular synthetic datasets via disjoint generative models. In this paradigm, a dataset is partitioned into disjoint subsets that are supplied to separate instances of generative models. The results are then combined post hoc by a joining operation that works in the absence of common variables/identifiers. The success of the framework is demonstrated through several case studies and examples on tabular data that help illuminate some of the design choices that one may make. The advantages achieved by the disjoint generation include: i) An observed increase in the empirical measurement of privacy. ii) Increased computational feasibility of certain model types. iii) Ability to generate synthetic data using a mixture of different generative models. Specifically, mixed-model synthesis bridges the gap between privacy and utility performance, providing highly competitive performance on Accuracy and Area Under the Curve for downstream tasks while significantly lowering the empirical re-identification risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。