arXiv:2502.14523cs.LGcs.CL2025-02被引 7

用GPT-4o零样本生成高质量表格数据,效果超越传统GAN模型。

Generative adversarial networks vs large language models: a comparative study on synthetic tabular data generation

  • 无需微调或真实数据预训练,仅靠自然语言提示生成表格数据。
  • 在均值、置信区间、相关性等指标上优于CTGAN,且隐私保护更强。
  • 适合需要快速生成高保真数据的研究者,尤其关注数据隐私与可解释性。

我们提出一种零样本生成合成表格数据的新框架。采用大语言模型GPT-4o与自然语言提示,无需任务特定微调或真实世界数据(RWD)预训练,即可生成高保真表格数据。为评估GPT-4o性能,我们在三个公开数据集(Iris、Fish Measurements、Real Estate Valuation)上将其与条件表格生成对抗网络(CTGAN)生成的数据在保真度和隐私性方面进行对比。尽管采用零样本方法,GPT-4o在保持均值、95%置信区间、双变量相关性及真实数据隐私方面仍优于CTGAN,且在样本量放大后表现更优。参数间的相关性方向与强度也得到一致保留。但分布特性仍需优化。结果表明,大语言模型在表格数据生成中具有潜力,可作为生成对抗网络与变分自编码器的便捷替代方案。

原文摘要 · Abstract (English)

We propose a new framework for zero-shot generation of synthetic tabular data. Using the large language model (LLM) GPT-4o and plain-language prompting, we demonstrate the ability to generate high-fidelity tabular data without task-specific fine-tuning or access to real-world data (RWD) for pre-training. To benchmark GPT-4o, we compared the fidelity and privacy of LLM-generated synthetic data against data generated with the conditional tabular generative adversarial network (CTGAN), across three open-access datasets: Iris, Fish Measurements, and Real Estate Valuation. Despite the zero-shot approach, GPT-4o outperformed CTGAN in preserving means, 95% confidence intervals, bivariate correlations, and data privacy of RWD, even at amplified sample sizes. Notably, correlations between parameters were consistently preserved with appropriate direction and strength. However, refinement is necessary to better retain distributional characteristics. These findings highlight the potential of LLMs in tabular data synthesis, offering an accessible alternative to generative adversarial networks and variational autoencoders.

表格生成大模型数据隐私零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。