arXiv:2503.05954cs.LG2025-03中稿 · Transactions on Ma…综述被引 14

综述表格数据生成的深度学习方法,帮用户选对工具

A Survey on Deep Learning Approaches for Tabular Data Generation: Utility, Alignment, Fidelity, Privacy, Diversity, and Beyond

  • 按实用、对齐、保真、隐私、多样性五类需求分类方法
  • 梳理各类模型的评估方式与相互关系,指明适用场景
  • 适合需要生成高质量合成数据的研究者和工程师

生成建模已成为合成表格数据的标准方法。然而,不同应用场景对合成数据的要求各异,以确保其实际可用性。本文从五个维度系统回顾了深度生成建模在表格数据生成中的应用:合成数据的实用性、与领域知识的一致性、统计分布与真实数据的保真度、隐私保护能力以及采样多样性。我们从两个层面组织现有方法:(i)按其所解决的需求分类,(ii)按底层模型架构分类。同时总结了每类需求的合适评估方法,分析了需求间的关联性,并梳理各类模型的特性。最后讨论了该领域的未来方向及评估方法的改进机会。整体上,本综述可作为表格数据生成的使用指南,帮助读者根据自身需求选择最合适的模型与评估手段。

原文摘要 · Abstract (English)

Generative modelling has become the standard approach for synthesising tabular data. However, different use cases demand synthetic data to comply with different requirements to be useful in practice. In this survey, we review deep generative modelling approaches for tabular data from the perspective of five types of requirements: utility of the synthetic data, alignment of the synthetic data with domain-specific knowledge, statistical fidelity of the synthetic data distribution compared to the real data distribution, privacy-preserving capabilities, and sampling diversity. We group the approaches along two levels of granularity: (i) based on the requirements they address and (ii) according to the underlying model they utilise. Additionally, we summarise the appropriate evaluation methods for each requirement, the relationships among the requirements, and the specific characteristics of each model type. Finally, we discuss future directions for the field, along with opportunities to improve the current evaluation methods. Overall, this survey can be seen as a user guide to tabular data generation: helping readers navigate available models and evaluation methods to find those best suited to their needs.

表格生成生成模型数据隐私评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。