arXiv:2409.01552cs.CLcs.AI2024-09ACL被引 11

通过自生成提示增强黑盒大模型的上下文学习能力。

Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs

  • 用自指导强化学习生成可靠衍生提示,构建上下文环境。
  • 在GPT-4等黑盒模型上显著提升响应质量,效果优于传统提示优化。
  • 适合希望提升黑盒模型性能的研究者与应用开发者。

大语言模型(LLMs)在生成高质量回复方面表现优异。为更好对齐人类偏好,已有诸多方法基于特定优化过程,但这些方法不适用于参数不可见的黑盒模型(如GPT-4)。黑盒模型性能高度依赖提示质量。现有提升方法常依赖提示优化模型,易导致优化后提示与原提示语义不一致,且忽略二者关系。为此,本文提出一种自指导上下文学习框架,使模型通过生成可靠的衍生提示来构建信息丰富的上下文环境。该方法结合自指导强化学习机制,在生成衍生提示时直接与响应模型交互,实现更好对齐。随后将查询任务建模为上下文学习问题,利用模型回复与衍生提示共同构成原始提示的上下文示范。该策略确保与原查询对齐,减少提示差异,最大化模型上下文学习能力。大量实验表明,该方法不仅能生成更可靠的衍生提示,还能显著提升模型响应有效性,包括在GPT-4等黑盒模型上。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown success in generating high-quality responses. In order to achieve better alignment with LLMs with human preference, various works are proposed based on specific optimization process, which, however, is not suitable to Black-Box LLMs like GPT-4, due to inaccessible parameters. In Black-Box LLMs case, their performance is highly dependent on the quality of the provided prompts. Existing methods to enhance response quality often involve a prompt refinement model, yet these approaches potentially suffer from semantic inconsistencies between the refined and original prompts, and typically overlook the relationship between them. To address these challenges, we introduce a self-instructed in-context learning framework that empowers LLMs to deliver more effective responses by generating reliable derived prompts to construct informative contextual environments. Our approach incorporates a self-instructed reinforcement learning mechanism, enabling direct interaction with the response model during derived prompt generation for better alignment. We then formulate querying as an in-context learning task, using responses from LLMs combined with the derived prompts to establish a contextual demonstration for the original prompt. This strategy ensures alignment with the original query, reduces discrepancies from refined prompts, and maximizes the LLMs' in-context learning capability. Extensive experiments demonstrate that the proposed method not only generates more reliable derived prompts but also significantly enhances LLMs' ability to deliver more effective responses, including Black-Box models such as GPT-4.

大模型提示工程黑盒模型上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。