用辩论和评分机制自动优化提示词,无需人工设定标准。
Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings
- 通过辩论+埃洛评分驱动提示词进化,结合优劣提示的优点。
- 在无真实答案的情况下,仍显著优于人工和现有方法。
- 适合需要主观评估的复杂任务,如创意生成、内容审核。
提示工程是制约大语言模型发挥潜力的关键瓶颈,尤其在涉及主观质量评估的任务中,难以定义明确的优化目标。现有自动化优化方法通常依赖特定任务的数值评分函数或通用模板,难以应对复杂场景。我们提出DEEVO(辩论驱动的提示词进化优化),通过辩论式评估与埃洛评分相结合的方式引导提示词演化。不同于以往方法,DEEVO在离散提示空间中探索,利用基于辩论反馈的智能交叉与策略性变异操作,融合成功与失败提示的优势,保持语义连贯性。采用埃洛评分作为适应度代理,同时促进性能提升并保留提示种群多样性。实验表明,在开放式与封闭式任务上,即使无真实标签反馈,DEEVO仍显著优于人工提示工程及现有先进方法。该框架将大模型推理能力与自适应优化结合,实现了无需预设指标的持续改进,推动了提示优化研究的发展。
原文摘要 · Abstract (English)
Prompt engineering represents a critical bottleneck to harness the full potential of Large Language Models (LLMs) for solving complex tasks, as it requires specialized expertise, significant trial-and-error, and manual intervention. This challenge is particularly pronounced for tasks involving subjective quality assessment, where defining explicit optimization objectives becomes fundamentally problematic. Existing automated prompt optimization methods falter in these scenarios, as they typically require well-defined task-specific numerical fitness functions or rely on generic templates that cannot capture the nuanced requirements of complex use cases. We introduce DEEVO (DEbate-driven EVOlutionary prompt optimization), a novel framework that guides prompt evolution through a debate-driven evaluation with an Elo-based selection. Contrary to prior work, DEEVOs approach enables exploration of the discrete prompt space while preserving semantic coherence through intelligent crossover and strategic mutation operations that incorporate debate-based feedback, combining elements from both successful and unsuccessful prompts based on identified strengths rather than arbitrary splicing. Using Elo ratings as a fitness proxy, DEEVO simultaneously drives improvement and preserves valuable diversity in the prompt population. Experimental results demonstrate that DEEVO significantly outperforms both manual prompt engineering and alternative state-of-the-art optimization approaches on open-ended tasks and close-ended tasks despite using no ground truth feedback. By connecting LLMs reasoning capabilities with adaptive optimization, DEEVO represents a significant advancement in prompt optimization research by eliminating the need of predetermined metrics to continuously improve AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。