用提示词让大模型精准生成带复杂依赖的表格数据
A Note on Statistically Accurate Tabular Data Generation Using Large Language Models
- 用概率驱动提示词估计条件分布
- 生成数据更符合真实统计特性
- 适合需要高保真表格数据的研究者
大型语言模型(LLMs)在合成表格数据方面展现出潜力,但现有方法难以保留特征间的复杂依赖关系,尤其是类别变量之间的依赖。本文提出一种概率驱动的提示方法,利用大模型估算条件分布,从而实现更准确、可扩展的数据合成。实验结果表明,通过提示概率分布能显著提升大模型生成表格数据的统计保真度。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise in synthetic tabular data generation, yet existing methods struggle to preserve complex feature dependencies, particularly among categorical variables. This work introduces a probability-driven prompting approach that leverages LLMs to estimate conditional distributions, enabling more accurate and scalable data synthesis. The results highlight the potential of prompting probability distributions to enhance the statistical fidelity of LLM-generated tabular data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。