arXiv:2604.23478cs.CL2026-04被引 6

测试大模型判卷时对提示语改写有多敏感,发现越大越不稳。

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

  • 构建手标提示对数据集,测不同表述下模型判断是否一致
  • 发现一致性与模型大小无关,最大新模型反而最不稳定
  • 共现位置偏差,尤其在成对比较任务中表现明显

大型语言模型被广泛用作自动化评估裁判,但其在语义等价提示重述下的判决稳定性尚未得到充分研究。我们对多个评估任务和裁判架构中的提示诱发决策不稳定性进行了系统性实证研究。为支持该分析,我们发布了JudgeSense基准,包含从主流NLP基准中提取的、经人工验证的提示-改写对,覆盖事实性、连贯性、相关性和偏好性,附带完整的决策日志。该基准可测量等价提示下的裁判稳定性,使研究者能评估稳定性是否与模型规模或指令微调相关,并识别对提示措辞最敏感的任务。评估显示,连贯性是区分裁判行为的主要任务,而事实性判断在标准条件下具有高稳定性。成对评估任务始终存在位置偏差。关键发现是:模型规模并非一致性的可靠代理;值得注意的是,最大最新模型并非最一致。

原文摘要 · Abstract (English)

Large language models are widely adopted as automated evaluation judges, yet the stability of their verdicts under semantically equivalent prompt rephrasings remains largely unexamined. We conduct a systematic empirical study of prompt-induced decision instability across multiple evaluation tasks and judge architectures. To facilitate this analysis, we release JudgeSense, a benchmark comprising hand-validated prompt-paraphrase pairs spanning factuality, coherence, relevance, and preference, drawn from established NLP benchmarks and accompanied by comprehensive decision logs. The benchmark enables the measurement of judge stability across equivalent prompts, allowing researchers to assess whether stability correlates with model scale or instruction-tuning, and to identify which tasks are most sensitive to prompt wording. Our evaluation reveals that coherence remains the primary task for distinguishing judge behavior, while factuality judgments demonstrate high stability under standard conditions. Pairwise evaluation tasks consistently exhibit position bias. Crucially, we find that model scale is not a reliable proxy for consistency; notably, as an interesting result in our analysis, the largest and newest models are not the most consistent.

大模型评测提示敏感性判卷稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。