评估金融合成数据的隐私与实用性权衡,发现不平衡数据下隐私保护更难。
Measuring Privacy Risks and Tradeoffs in Financial Synthetic Data Generation
- 对比多种生成模型在金融表格数据上的表现
- 发现类别严重不均衡时隐私保护能力显著下降
- 适合关注数据合规与隐私风险的研究者
我们研究了在表格型金融数据集上合成数据生成方法的隐私-效用权衡,该领域具有高监管风险和严重的类别不平衡问题。考虑了代表性的表格数据生成器,包括自编码器、生成对抗网络、扩散模型和拷贝拉合成器。为应对金融领域的挑战,我们提出了 GAN 和自编码器合成器的新隐私保护实现。通过对比平衡与不平衡输入数据集,评估生成器在数据质量、下游任务效用和隐私保护方面的综合表现。结果揭示了在存在严重类别不平衡和混合类型属性的数据中生成合成数据的独特挑战。
原文摘要 · Abstract (English)
We explore the privacy-utility tradeoff of synthetic data generation schemes on tabular financial datasets, a domain characterized by high regulatory risk and severe class imbalance. We consider representative tabular data generators, including autoencoders, generative adversarial networks, diffusion, and copula synthesizers. To address the challenges of the financial domain, we provide novel privacy-preserving implementations of GAN and autoencoder synthesizers. We evaluate whether and how well the generators simultaneously achieve data quality, downstream utility, and privacy, with comparison across balanced and imbalanced input datasets. Our results offer insight into the distinct challenges of generating synthetic data from datasets that exhibit severe class imbalance and mixed-type attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。