分析合成数据在共享与增强中的适用性,揭示其局限性。
Should I use Synthetic Data for That? An Analysis of the Suitability of Synthetic Data for Data Sharing and Augmentation
- 从隐私保护、模型训练、统计估计三方面建模分析
- 发现合成数据在多数场景下效果有限,存在根本性约束
- 为决策者提供判断合成数据是否适用的评估框架
生成模型的进展使合成数据被视为解决数据访问、稀缺和代表性不足问题的首选方案。本文研究三个典型应用场景:(1)以合成数据替代专有数据集进行统计分析并保护隐私;(2)用合成数据扩充机器学习训练集以提升性能;(3)通过合成数据减少统计估计的方差。针对每种场景,我们形式化问题设定,并结合形式化分析与案例研究,探讨合成数据实现目标的条件。分析揭示了其在实际应用中存在根本性与实践性限制,许多现有或设想的应用场景并不适合使用合成数据。我们的形式化框架与分类体系,可帮助决策者评估合成数据是否适配其具体数据问题。
原文摘要 · Abstract (English)
Recent advances in generative modelling have led many to see synthetic data as the go-to solution for a range of problems around data access, scarcity, and under-representation. In this paper, we study three prominent use cases: (1) Sharing synthetic data as a proxy for proprietary datasets to enable statistical analyses while protecting privacy, (2) Augmenting machine learning training sets with synthetic data to improve model performance, and (3) Augmenting datasets with synthetic data to reduce variance in statistical estimation. For each use case, we formalise the problem setting and study, through formal analysis and case studies, under which conditions synthetic data can achieve its intended objectives. We identify fundamental and practical limits that constrain when synthetic data can serve as an effective solution for a particular problem. Our analysis reveals that due to these limits many existing or envisioned use cases of synthetic data are a poor problem fit. Our formalisations and classification of synthetic data use cases enable decision makers to assess whether synthetic data is a suitable approach for their specific data availability problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。