arXiv:2507.15675cs.CL2025-07ACL被引 3

同时优化系统与用户提示,提升大模型推理表现

P3: Prompts Promote Prompting

  • 双提示协同迭代优化,突破单一提示改进局限
  • 在Arena-hard、GSM8K等任务上性能显著超越基线
  • 适合需要高效提示工程的AI研发与应用者

当前大语言模型应用常采用包含系统提示和用户提示的多组件提示来引导模型行为。尽管近期研究已证明自动优化系统或用户提示可提升性能,但因两组件相互依赖,单方面优化常导致次优结果。本文提出P3框架,通过迭代过程同步优化系统与用户提示。离线优化后的提示进一步用于在线提示中,实现基于查询的动态优化。在通用任务(如Arena-hard、Alpaca-eval)和推理任务(如GSM8K、GPQA)上的大量实验表明,P3在自动提示优化领域表现更优。结果验证了整体优化策略在多领域提升大模型性能的有效性。

原文摘要 · Abstract (English)

Current large language model (LLM) applications often employ multi-component prompts, comprising both system and user prompts, to guide model behaviors. While recent advancements have demonstrated the efficacy of automatically optimizing either the system or user prompt to boost performance, such unilateral approaches often yield suboptimal outcomes due to the interdependent nature of these components. In this work, we introduce P3, a novel self-improvement framework that concurrently optimizes both system and user prompts through an iterative process. The offline optimized prompts are further leveraged to promote online prompting by performing query-dependent prompt optimization. Extensive experiments on general tasks (e.g., Arena-hard and Alpaca-eval) and reasoning tasks (e.g., GSM8K and GPQA) demonstrate that P3 achieves superior performance in the realm of automatic prompt optimization. Our results highlight the effectiveness of a holistic optimization strategy in enhancing LLM performance across diverse domains.

提示工程大模型优化框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。