用进化算法系统探索提示词空间,发现结构差异如何影响大模型表现。
Diverse Prompts: Illuminating the Prompt Space of Large Language Models with MAP-Elites
- 结合无上下文语法与MAP-Elites算法,自动生成高质量多样提示词。
- 在7个BigBench Lite任务中验证,提示词结构显著影响模型性能。
- 适合需要定制化提示词的开发者和研究者参考使用。
提示工程对优化大语言模型至关重要,但提示结构与任务表现之间的关系仍不明确。本文提出一种融合无上下文语法(CFG)与MAP-Elites算法的进化方法,系统探索提示词空间。该方法兼顾质量与多样性,生成高性能且结构多样的提示词,并通过改变示例数量(shots)和推理深度等特征,分析其在多种任务中的适配性。通过系统映射表型空间,揭示结构变化对大模型性能的影响,为任务特定及可适应的提示设计提供可操作洞见。在多个大模型上评估了7个BigBench Lite任务,结果表明质量与多样性之间存在关键协同作用,显著提升了大模型的有效性与通用性。
原文摘要 · Abstract (English)
Prompt engineering is essential for optimizing large language models (LLMs), yet the link between prompt structures and task performance remains underexplored. This work introduces an evolutionary approach that combines context-free grammar (CFG) with the MAP-Elites algorithm to systematically explore the prompt space. Our method prioritizes quality and diversity, generating high-performing and structurally varied prompts while analyzing their alignment with diverse tasks by varying traits such as the number of examples (shots) and reasoning depth. By systematically mapping the phenotypic space, we reveal how structural variations influence LLM performance, offering actionable insights for task-specific and adaptable prompt design. Evaluated on seven BigBench Lite tasks across multiple LLMs, our results underscore the critical interplay of quality and diversity, advancing the effectiveness and versatility of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。