对比现代表格数据生成技术,为隐私保护与实用性的平衡提供指南。
Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques
- 按下游应用、隐私保障和数据效用设计新分类体系
- 提出基准框架,推动技术与真实需求对齐
- 适合关注隐私数据生成的科研与工程人员
随着隐私法规日益严格,真实数据获取愈发受限,合成数据生成已成为重要解决方案,尤其在金融、医疗和社科等依赖表格数据的领域。本综述系统梳理了近期合成表格数据生成的进展,重点关注保持复杂特征关系、统计保真度及满足隐私要求的方法。核心贡献是提出基于实际生成目标的新分类体系,涵盖下游应用、隐私保证和数据效用,直接指导方法设计与评估策略。同时强调条件生成与风险敏感建模等可操作目标。此外,本文构建基准框架,促进技术创新与真实需求对接。该工作兼具理论深度与实践价值,为未来研究提供路线图,并指导隐私敏感环境中的合成数据部署。
原文摘要 · Abstract (English)
As privacy regulations become more stringent and access to real-world data becomes increasingly constrained, synthetic data generation has emerged as a vital solution, especially for tabular datasets, which are central to domains like finance, healthcare and the social sciences. This survey presents a comprehensive and focused review of recent advances in synthetic tabular data generation, emphasizing methods that preserve complex feature relationships, maintain statistical fidelity, and satisfy privacy requirements. A key contribution of this work is the introduction of a novel taxonomy based on practical generation objectives, including intended downstream applications, privacy guarantees, and data utility, directly informing methodological design and evaluation strategies. Therefore, this review prioritizes the actionable goals that drive synthetic data creation, including conditional generation and risk-sensitive modeling. Additionally, the survey proposes a benchmark framework to align technical innovation with real-world demands. By bridging theoretical foundations with practical deployment, this work serves as both a roadmap for future research and a guide for implementing synthetic tabular data in privacy-critical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。