用知识代替大量示例,让提示词更高效生成高质量合成数据。
The Prompt is Mightier than the Example
- 将领域知识显式注入提示词,减少对示例的依赖。
- 实验证明知识越多、示例越少,合成数据质量仍可保持较高水平。
- 适合需要低成本生成数据的研究者或工业应用。
近年来,链式思维等提示优化方法显著提升了大语言模型(LLM)生成内容的质量。上下文学习(ICL)通过少量代表性示例引导生成,也在合成表格数据生成中表现优异,能近似复杂异构分布的数据样本。然而,保证高保真度合成数据通常需要大量ICL示例,而这些示例可能难以获取或成本高昂。随着大模型规模扩大,其内置先验知识日益丰富,可替代特定数据示例。本文提出知识引导提示(KGP),探索其在提示优化中替代ICL示例的能力。我们研究‘提示能替代多少示例?’这一问题,通过显式注入领域知识(推断或已知)到提示中,降低对ICL的依赖。系统实验揭示了ICL与KGP之间的权衡关系,发现生成数据质量随领域知识增加、示例数量减少而呈现可量化的经验缩放规律。结果表明,KGP可作为可扩展的替代或补充方案,为合成数据生成开辟新路径。
原文摘要 · Abstract (English)
Numerous recent prompt optimization approaches like chain-of-thought, have been demonstrated to significantly improve the quality of content generated by large language models (LLMs). In-context learning (ICL), a recent paradigm where a few representative examples guide content generation has also led to strong improvements in generation quality of LLM generated content. This idea has been applied to great effect in synthetic tabular data generation, where LLMs, through effective use of ICL and prompt optimization, can generate data that approximate samples from complex, heterogeneous distributions based on representative examples. However, ensuring high-fidelity synthetic data often requires a very large number of ICL examples which may be unavailable or costly to obtain. At the same time, as LLMs get larger and larger, their in-built prior knowledge becomes vast and can potentially substitute for specific data examples. In this paper, we introduce Knowledge-Guided Prompting (KGP) as a new knob in prompt optimization and explore the ability of KGP-based prompt optimization to offset the cost of ICL. Specifically, we explore the question `how many examples can a prompt substitute for?' and explore knowledge-guided prompting (KGP) where domain knowledge, either inferred or available, is explicitly injected into the prompt, reducing dependence on ICL examples. Our experiments systematically explore the trade-off between ICL and KGP, revealing an empirical scaling law that quantifies how quality of generated synthetic data varies with increasing domain knowledge and decreasing example count. Our results demonstrate that knowledge-guided prompting can be a scalable alternative, or addition, to in-context examples, unlocking new approaches to synthetic data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。