arXiv:2509.04152cs.LGcs.AI2025-09被引 6

用智能体流程生成高质量表格数据,无需训练大模型即可达到顶尖效果。

TAGAL: Tabular Data Generation using Agentic LLM Methods

  • 采用智能体工作流,通过反馈迭代提升合成数据质量。
  • 在多种数据集上表现媲美需训练大模型的先进方法。
  • 适合需要快速生成真实数据替代品的研究者与工程师。

数据生成是提升机器学习任务性能的常用方法,尤其在分类模型训练中。本文提出TAGAL,一种基于智能体工作流的合成表格数据生成方法。该方法利用大语言模型(LLM)实现自动、迭代的数据生成过程,通过反馈机制优化数据质量,且无需额外训练大模型。大语言模型的应用还支持在生成过程中引入外部知识。我们在多个数据集上评估了TAGAL在生成数据质量方面的表现,包括下游机器学习模型的实用性——既在仅使用合成数据训练分类器时,也结合真实与合成数据进行测试。此外,我们对比了真实数据与生成数据之间的相似性。结果表明,TAGAL在性能上可与需要训练大模型的前沿方法比肩,通常优于其他无需训练的方法。这些发现凸显了智能体工作流在大模型驱动数据生成中的潜力,并为未来研究开辟新方向。

原文摘要 · Abstract (English)

The generation of data is a common approach to improve the performance of machine learning tasks, among which is the training of models for classification. In this paper, we present TAGAL, a collection of methods able to generate synthetic tabular data using an agentic workflow. The methods leverage Large Language Models (LLMs) for an automatic and iterative process that uses feedback to improve the generated data without any further LLM training. The use of LLMs also allows for the addition of external knowledge in the generation process. We evaluate TAGAL across diverse datasets and different aspects of quality for the generated data. We look at the utility of downstream ML models, both by training classifiers on synthetic data only and by combining real and synthetic data. Moreover, we compare the similarities between the real and the generated data. We show that TAGAL is able to perform on par with state-of-the-art approaches that require LLM training and generally outperforms other training-free approaches. These findings highlight the potential of agentic workflow and open new directions for LLM-based data generation methods.

数据生成大模型表格数据智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。