用序贯优化自动设计高效提示词,节省评估成本
A Sequential Optimal Learning Approach to Automated Prompt Engineering in Large Language Models
- 基于特征表示提示词,扩大搜索空间
- 采用贝叶斯回归与前瞻知识梯度策略,提升学习效率
- 适合评估成本高的场景,如指令归纳任务
设计有效的提示词对引导大语言模型生成期望输出至关重要。自动化提示工程旨在减少人工依赖,实现提示词的设计、优化与迭代的自动化。本文提出一种面向自动化提示工程的最优学习框架,通过序贯方式识别有效提示特征,并在有限评估预算下高效分配资源。我们引入基于特征的提示表达方法,显著扩展搜索空间;利用贝叶斯回归捕捉相似提示间的相关性,加速学习过程。为高效探索大规模提示特征空间以获得高质量提示,采用前向知识梯度(KG)策略进行序贯最优学习。KG策略通过求解混合整数二次锥优化问题实现高效计算,具备可扩展性,且适用于仅由约束定义的提示。在指令归纳任务上的实验表明,该方法显著优于多个基准策略,验证了在有限评估预算下使用KG策略的优越性。本框架为高成本提示评估场景下的自动化提示工程部署提供了可行方案。
原文摘要 · Abstract (English)
Designing effective prompts is essential to guiding large language models (LLMs) toward desired responses. Automated prompt engineering aims to reduce reliance on manual effort by streamlining the design, refinement, and optimization of natural language prompts. This paper proposes an optimal learning framework for automated prompt engineering, designed to sequentially identify effective prompt features while efficiently allocating a limited evaluation budget. We introduce a feature-based method to express prompts, which significantly broadens the search space. Bayesian regression is employed to utilize correlations among similar prompts, accelerating the learning process. To efficiently explore the large space of prompt features for a high quality prompt, we adopt the forward-looking Knowledge-Gradient (KG) policy for sequential optimal learning. The KG policy is computed efficiently by solving mixed-integer second-order cone optimization problems, making it scalable and capable of accommodating prompts characterized only through constraints. We demonstrate that our method significantly outperforms a set of benchmark strategies assessed on instruction induction tasks. The results highlight the advantages of using the KG policy for prompt learning given a limited evaluation budget. Our framework provides a solution to deploying automated prompt engineering in a wider range applications where prompt evaluation is costly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。