提出统一评估合成表格数据的框架FEST,兼顾隐私与实用性的平衡。
FEST: A Unified Framework for Evaluating Synthetic Tabular Data
- 整合攻击型与距离型隐私度量,结合相似性与机器学习效用指标
- 在多个数据集上验证了不同生成模型的隐私-效用权衡表现
- 开源工具包支持研究者评估合成数据质量,适合隐私保护研究者使用
合成数据生成利用生成式机器学习技术,为缓解真实数据使用中的隐私问题提供了前景。合成数据在保持强隐私保障的同时,能高度模仿真实数据。然而,当前仍缺乏对合成数据生成的全面评估框架,尤其在隐私保护与数据效用之间的平衡方面。本研究提出FEST,一个系统化的合成表格数据评估框架。FEST融合多种隐私度量(基于攻击和基于距离),以及相似性和机器学习效用度量,实现全面评估。我们开发了基于Python的开源FEST库,并在多个数据集上验证其有效性,展示了其在分析不同合成数据生成模型隐私-效用权衡方面的性能。FEST源代码已发布于Github。
原文摘要 · Abstract (English)
Synthetic data generation, leveraging generative machine learning techniques, offers a promising approach to mitigating privacy concerns associated with real-world data usage. Synthetic data closely resembles real-world data while maintaining strong privacy guarantees. However, a comprehensive assessment framework is still missing in the evaluation of synthetic data generation, especially when considering the balance between privacy preservation and data utility in synthetic data. This research bridges this gap by proposing FEST, a systematic framework for evaluating synthetic tabular data. FEST integrates diverse privacy metrics (attack-based and distance-based), along with similarity and machine learning utility metrics, to provide a holistic assessment. We develop FEST as an open-source Python-based library and validate it on multiple datasets, demonstrating its effectiveness in analyzing the privacy-utility trade-off of different synthetic data generation models. The source code of FEST is available on Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。