系统梳理生成合成表格数据的主流方法与挑战,覆盖生成、评估到应用全链路。
A Comprehensive Survey of Synthetic Tabular Data Generation
- 按生成范式分类:传统方法、扩散模型、大语言模型三类并行比较
- 整合最新进展,涵盖生成质量、适用场景与隐私保护能力对比
- 适合想快速掌握该领域全貌的研究者与工业界落地人员
表格数据在医疗、金融、教育等现实应用中极为常见,但其机器学习应用常受限于数据稀缺、隐私问题及类别不平衡。合成表格数据生成通过生成模型学习真实数据分布,生成兼具真实性与隐私保护性的样本,成为重要解决方案。尽管该领域关注度持续上升,现有综述多聚焦单一方法(如GAN或隐私增强技术),缺乏对扩散模型与大语言模型等新进展的统一整合。本文系统梳理合成表格数据生成方法,分为三大核心部分:背景部分介绍生成流程,包括问题定义、生成方法、后处理与评估;方法部分将现有技术分为传统生成方法、扩散模型方法和基于大语言模型的方法,从架构、生成质量与适用性进行对比;应用与挑战部分总结实际应用场景、常用数据集,并讨论异构性、数据保真度与隐私保护等开放问题。本综述旨在为研究者与实践者提供全景式理解,并指明未来关键方向。
原文摘要 · Abstract (English)
Tabular data is one of the most prevalent and important data formats in real-world applications such as healthcare, finance, and education. However, its effective use in machine learning is often constrained by data scarcity, privacy concerns, and class imbalance. Synthetic tabular data generation has emerged as a powerful solution, leveraging generative models to learn underlying data distributions and produce realistic, privacy-preserving samples. Although this area has seen growing attention, most existing surveys focus narrowly on specific methods (e.g., GANs or privacy-enhancing techniques), lacking a unified and comprehensive view that integrates recent advances such as diffusion models and large language models (LLMs). In this survey, we present a structured and in-depth review of synthetic tabular data generation methods. Specifically, the survey is organized into three core components: (1) Background, which covers the overall generation pipeline, including problem definitions, synthetic tabular data generation methods, post processing, and evaluation; (2) Generation Methods, where we categorize existing approaches into traditional generation methods, diffusion model methods, and LLM-based methods, and compare them in terms of architecture, generation quality, and applicability; and (3) Applications and Challenges, which summarizes practical use cases, highlights common datasets, and discusses open challenges such as heterogeneity, data fidelity, and privacy protection. This survey aims to provide researchers and practitioners with a holistic understanding of the field and to highlight key directions for future work in synthetic tabular data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。