arXiv:2603.10477cs.CL2026-03中稿 · IEEE Access

为提示词和回复提供可解释的联合评估,助力优化大模型交互效果。

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses

  • 构建9维度评分框架,融合提示清晰度与回复质量,支持多维度分析。
  • 在7个基准上验证,与传统准确率高度一致(Spearman相关0.97)。
  • 生成自然语言理由,帮助诊断问题,适合提示工程优化与教学使用。

提示设计是控制大语言模型的核心方式,但现有评估通常仅关注答案正确性,难以解释成败原因且缺乏可操作指导。本文提出PEEM(提示工程评估指标),一个统一的、可解释的提示与回复联合评估框架。PEEM定义包含3个提示标准(清晰度/结构、语言质量、公平性)和6个回复标准(准确性、连贯性、相关性、客观性、清晰度、简洁性)的结构化评分体系,并使用大模型评估器输出(1)1-5分制的标量得分,(2)基于评分维度的自然语言解释。在7个基准和5个任务模型上,PEEM的准确性维度与传统准确率高度一致(总体斯皮尔曼相关约0.97,皮尔逊相关约0.94,p < 0.001)。四模型多评估者研究显示判断结果稳定(成对相关系数0.68-0.85),支持跨评估器部署。此外,PEEM能捕捉互补的语言失败模式,在提示扰动下表现稳健:提示质量趋势与下游准确率变化同步,语义对抗扰动导致评分明显下降,语义保持的改写则具有高稳定性(鲁棒率约76.7%-80.6%)。仅依赖PEEM评分与解释作为反馈,零样本提示重写循环可使下游准确率提升最高达11.7点,优于监督与强化学习基线。总体而言,PEEM提供可复现、以准则为导向的协议,连接提示设计与响应行为,支持系统性诊断与优化大模型交互。

原文摘要 · Abstract (English)

Prompt design is a primary control interface for large language models (LLMs), yet standard evaluations largely reduce performance to answer correctness, obscuring why a prompt succeeds or fails and providing little actionable guidance. We propose PEEM (Prompt Engineering Evaluation Metrics), a unified framework for joint and interpretable evaluation of both prompts and responses. PEEM defines a structured rubric with 9 axes: 3 prompt criteria (clarity/structure, linguistic quality, fairness) and 6 response criteria (accuracy, coherence, relevance, objectivity, clarity, conciseness), and uses an LLM-based evaluator to output (i) scalar scores on a 1-5 Likert scale and (ii) criterion-specific natural-language rationales grounded in the rubric. Across 7 benchmarks and 5 task models, PEEM's accuracy axis strongly aligns with conventional accuracy while preserving model rankings (aggregate Spearman rho about 0.97, Pearson r about 0.94, p < 0.001). A multi-evaluator study with four models shows consistent relative judgments (pairwise rho = 0.68-0.85), supporting evaluator-agnostic deployment. Beyond alignment, PEEM captures complementary linguistic failure modes and remains informative under prompt perturbations: prompt-quality trends track downstream accuracy under iterative rewrites, semantic adversarial manipulations induce clear score degradation, and meaning-preserving paraphrases yield high stability (robustness rate about 76.7-80.6%). Finally, using only PEEM scores and rationales as feedback, a zero-shot prompt rewriting loop improves downstream accuracy by up to 11.7 points, outperforming supervised and RL-based prompt-optimization baselines. Overall, PEEM provides a reproducible, criterion-driven protocol that links prompt formulation to response behavior and enables systematic diagnosis and optimization of LLM interactions.

提示工程评估框架可解释性LLM优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。