arXiv:2509.01790cs.CLcs.AI2025-09EMNLP被引 24

重新审视大模型提示敏感性,发现多数问题源于评估方法。

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs

  • 用大模型自评代替传统打分,减少评估偏差。
  • 不同提示下模型表现差异大幅降低,排名相关性显著提升。
  • 适合关注评测公平性与模型真实能力的研究者。

提示敏感性指同一内容用不同表述(如同义词或改写)会导致大语言模型(LLM)性能显著波动,被广泛认为是其核心缺陷。本文系统评估7个主流LLM(如GPT、Gemini系列)在6个基准上的表现,涵盖12种不同提示模板的多项选择与开放问答任务。结果发现,多数敏感性现象源于评估方法的局限,包括基于对数似然的打分和严格答案匹配,这些方法常忽略语义正确但表达不同的回答。当采用‘大模型作为裁判’(LLM-as-a-Judge)的评估方式时,性能方差明显下降,模型排名在不同提示间的相关性显著提高。研究表明,现代大模型对提示形式的鲁棒性远超以往认知,提示敏感性更多是评估过程的人为产物,而非模型本质缺陷。

原文摘要 · Abstract (English)

Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large language model (LLM) performance, has been widely accepted as a core limitation of LLMs. In this work, we revisit this issue and ask: Is the widely reported high prompt sensitivity truly an inherent weakness of LLMs, or is it largely an artifact of evaluation processes? To answer this question, we systematically evaluate 7 LLMs (e.g., GPT and Gemini family) across 6 benchmarks, including both multiple-choice and open-ended tasks on 12 diverse prompt templates. We find that much of the prompt sensitivity stems from heuristic evaluation methods, including log-likelihood scoring and rigid answer matching, which often overlook semantically correct responses expressed through alternative phrasings, such as synonyms or paraphrases. When we adopt LLM-as-a-Judge evaluations, we observe a substantial reduction in performance variance and a consistently higher correlation in model rankings across prompts. Our findings suggest that modern LLMs are more robust to prompt templates than previously believed, and that prompt sensitivity may be more an artifact of evaluation than a flaw in the models.

大模型评测提示工程评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。