对比六种关系型数据生成方法,发现无一能完全还原原始数据。
Benchmarking the Fidelity and Utility of Synthetic Relational Data
- 构建新评测工具,融合最佳实践与鲁棒检测方法
- 生成数据与真实数据仍可区分,预测性能相关性中等
- 适合关注数据隐私与合成质量评估的研究者
关系型数据合成因表间关联复杂而更具挑战性,现有评测方法尚不完善。本文系统梳理相关工作、常用数据集及评估指标,整合最佳实践并提出新型鲁棒检测方法,构建基准评测工具,对六种合成方法(含两种商用工具)进行对比。结果表明,现有方法均无法生成与原始数据难以区分的合成数据;在实用性方面,合成数据与真实数据在模型预测性能和特征重要性上的相关性多为中等水平。
原文摘要 · Abstract (English)
Synthesizing relational data has started to receive more attention from researchers, practitioners, and industry. The task is more difficult than synthesizing a single table due to the added complexity of relationships between tables. For the same reason, benchmarking methods for synthesizing relational data introduces new challenges. Our work is motivated by a lack of an empirical evaluation of state-of-the-art methods and by gaps in the understanding of how such an evaluation should be done. We review related work on relational data synthesis, common benchmarking datasets, and approaches to measuring the fidelity and utility of synthetic data. We combine the best practices and a novel robust detection approach into a benchmarking tool and use it to compare six methods, including two commercial tools. While some methods are better than others, no method is able to synthesize a dataset that is indistinguishable from original data. For utility, we typically observe moderate correlation between real and synthetic data for both model predictive performance and feature importance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。