arXiv:2412.18111cs.AI2024-12被引 7

用提示词增强生成高质量合成表格数据,突破大模型长度限制

AIGT: AI Generative Table Based on Prompt

  • 利用表描述、模式等元信息做提示词,提升生成质量
  • 在20个公开数据集中的14个和两个工业数据集上达到领先性能
  • 通过长文本分块算法克服大模型令牌限制,可处理任意规模表格

表格数据占企业数据资产的80%以上,在多个领域至关重要。随着隐私保护和数据共享限制日益严格,生成高质量合成表格数据变得尤为关键。近年来,大型语言模型(LLMs)通过利用语义信息,有效生成真实感强的表格数据,克服了因独热编码带来的高维难题。然而,现有方法未能充分挖掘表格中丰富的信息。为此,我们提出基于提示增强的AI生成表格(AIGT),利用表描述、模式等元信息作为提示,生成超高质量合成数据。为解决LLM的令牌长度限制,我们设计了长文本分块算法,使AIGT可建模任意规模的表格。AIGT在20个公共数据集中的14个以及两个支付宝风控系统的真实工业数据集上表现领先。

原文摘要 · Abstract (English)

Tabular data, which accounts for over 80% of enterprise data assets, is vital in various fields. With growing concerns about privacy protection and data-sharing restrictions, generating high-quality synthetic tabular data has become essential. Recent advancements show that large language models (LLMs) can effectively gener-ate realistic tabular data by leveraging semantic information and overcoming the challenges of high-dimensional data that arise from one-hot encoding. However, current methods do not fully utilize the rich information available in tables. To address this, we introduce AI Generative Table (AIGT) based on prompt enhancement, a novel approach that utilizes meta data information, such as table descriptions and schemas, as prompts to generate ultra-high quality synthetic data. To overcome the token limit constraints of LLMs, we propose long-token partitioning algorithms that enable AIGT to model tables of any scale. AIGT achieves state-of-the-art performance on 14 out of 20 public datasets and two real industry datasets within the Alipay risk control system.

表格生成大模型合成数据提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。