arXiv:2409.03946cs.CL2024-09中稿 · IEEE ICASSP 2025被引 2

用领域知识增强提示词,提升大模型生成表格数据的质量与效率

On The Role of Prompt Construction In Enhancing Efficacy and Efficiency of LLM-Based Tabular Data Generation

  • 通过专家、大模型和新映射三种方式构建带上下文的提示词
  • 上下文丰富的提示词使生成数据质量显著提升,训练效率提高
  • 适合需要高质量合成表格数据的研究者或工业应用

基于大模型的真实世界表格数据生成常因列名缺乏充分语义上下文而受阻。我们假设,在提示词中融入领域特定知识可同时提升生成质量和效率。为验证该假设,我们探索了三种提示构造策略:专家引导、大模型引导和新型映射。在最近提出的GReaT框架上进行的实证研究发现,富含上下文的提示词能显著提升数据生成质量与训练效率。

原文摘要 · Abstract (English)

LLM-based data generation for real-world tabular data can be challenged by the lack of sufficient semantic context in feature names used to describe columns. We hypothesize that enriching prompts with domain-specific insights can improve both the quality and efficiency of data generation. To test this hypothesis, we explore three prompt construction protocols: Expert-guided, LLM-guided, and Novel-Mapping. Through empirical studies with the recently proposed GReaT framework, we find that context-enriched prompts lead to significantly improved data generation quality and training efficiency.

表格生成提示工程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。