提出可指导优化的可靠提示评估信号,让提示修改更精准高效。
Knowing How to Edit: Reliable Evaluation Signals for Diagnosing and Optimizing Prompts at Query Level
- 构建多指标融合的评估框架,无需执行模型即可预测提示质量。
- 评估准确率达83.7%,显著优于传统方法。
- 适合需要可解释性提示优化的研究者和工程落地场景。
提示优化已成为激发大语言模型强性能的核心机制,近期研究提出了多种提示评估指标与优化策略。然而,评估与优化常被割裂,限制了评估对优化的指导作用。本文将提示优化视为由性能相关评估信号驱动的过程,提出一种评估引导的提示优化方法,显式连接评估与查询相关的优化。该方法将多个互补的提示质量指标整合进一个反映性能的评估框架,并训练一个无需执行模型的评估器,直接从文本预测提示质量,避免重复调用模型。这些评估信号以目标明确且可解释的方式指导提示改进。实验表明,所提评估器在预测提示性能上达到83.7%的准确率。将其引入优化流程后,本方法在八个基准数据集和三种不同主干LLM上均持续优于现有基线。结果表明,可靠高效的评估信号可成为鲁棒、可解释提示优化的坚实基础。
原文摘要 · Abstract (English)
Prompt optimization has become a central mechanism for eliciting strong performance from LLMs, and recent work has made substantial progress by proposing diverse prompt evaluation metrics and optimization strategies. Despite these advances, prompt evaluation and prompt optimization are often developed in isolation, limiting the extent to which evaluation can effectively inform prompt refinement. In this work, we study prompt optimization as a process guided by performance-relevant evaluation signals. To address the disconnect between evaluation and optimization, we propose an evaluation-instructed prompt optimization approach that explicitly connects prompt evaluation with query-dependent optimization. Our method integrates multiple complementary prompt quality metrics into a performance-reflective evaluation framework and trains an execution-free evaluator that predicts prompt quality directly from text, avoiding repeated model executions. These evaluation signals then guide prompt refinement in a targeted and interpretable manner. Empirically, the proposed evaluator achieves 83.7% accuracy in predicting prompt performance. When incorporated into the optimization process, our approach consistently outperforms existing optimization baselines across eight benchmark datasets and three different backbone LLMs. Overall, our results demonstrate that reliable and efficient evaluation signals can serve as an effective foundation for robust and interpretable prompt optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。