优化系统提示词可让大模型通用表现媲美专用提示。
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
- 用遗传算法迭代拼装提示组件,自动优化系统提示。
- 一个优化后的系统提示在47类任务上表现媲美定制任务提示。
- 效果跨模型、语言和参数量均稳定,适合通用场景部署。
大型语言模型在多种场景中表现出色,但其性能部分依赖于提示的选择。以往研究多聚焦于特定任务的提示优化,却较少关注提示中通用指令部分——即系统提示。为此,我们提出SPRIG,一种基于编辑的遗传算法,通过预设组件迭代构建提示,以提升模型在通用场景下的表现。我们在47种不同类型的任务上评估系统提示性能,验证其泛化能力。结果表明,单一优化后的系统提示表现可与为每个任务单独优化的提示相媲美。进一步地,结合系统提示与任务提示优化能带来额外提升,凸显二者互补性。实验还发现,优化后的系统提示在不同模型家族、参数规模及语言间具有良好泛化能力。本研究揭示了系统级指令在释放大模型潜力中的关键作用。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown impressive capabilities in many scenarios, but their performance depends, in part, on the choice of prompt. Past research has focused on optimizing prompts specific to a task. However, much less attention has been given to optimizing the general instructions included in a prompt, known as a system prompt. To address this gap, we propose SPRIG, an edit-based genetic algorithm that iteratively constructs prompts from prespecified components to maximize the model's performance in general scenarios. We evaluate the performance of system prompts on a collection of 47 different types of tasks to ensure generalizability. Our study finds that a single optimized system prompt performs on par with task prompts optimized for each individual task. Moreover, combining system and task-level optimizations leads to further improvement, which showcases their complementary nature. Experiments also reveal that the optimized system prompts generalize effectively across model families, parameter sizes, and languages. This study provides insights into the role of system-level instructions in maximizing LLM potential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。