arXiv:2603.11027cs.CL2026-03被引 3

大模型评分看似一致,实则依赖表面规则;用知识驱动的评分标准可提升评估可信度。

Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge

  • 发现评分一致性多源于共享表面规则,而非真实质量判断
  • 高质输出反而评分最不一致,共识常为虚假现象
  • 基于领域知识动态生成评分标准,显著提升评估可靠性

LLM作为评分者依赖于高评价者一致性的假设。我们提出两项互补发现:第一,这种一致性常为虚假。通过105,600个评估实例的大规模研究(32个LLM × 3个前沿评委 × 100个任务 × 11种温度),发现模型级一致性极高(Spearman ρ=0.99),但样本级一致性仅中等(Pearson 平均r=0.72;绝对一致性ICC=0.67),且仅靠共享评分框架就能恢复62%的一致性。高质量输出反而获得最不一致的评分。第二,基于领域知识动态生成评分标准可提升评估意义。我们提出MERG(元认知增强的评分框架),其在教育(+22%)、学术(+27%)等有明确标准的领域提升了一致性,而在主观性强的领域则因真实分歧而下降。这表明应动态融合专家知识,而非依赖通用标准,对RLAIF中的奖励建模有重要启示。

原文摘要 · Abstract (English)

The paradigm of LLM-as-a-judge relies on a critical assumption, namely that high inter-evaluator agreement indicates reliable and objective evaluation. We present two complementary findings that challenge this assumption. \textbf{First}, we demonstrate that this consensus is frequently illusory. We identify and formalize \textbf{Evaluation Illusion}, a phenomenon where LLM judges generate sophisticated critiques yet anchor scores on shared surface heuristics rather than substantive quality. Through a large-scale study of 105,600 evaluation instances (32 LLMs $\times$ 3 frontier judges $\times$ 100 tasks $\times$ 11 temperatures), we show that model-level agreement (Spearman $ρ= 0.99$) masks fragile sample-level agreement (Pearson $\bar{r} = 0.72$; absolute agreement ICC $= 0.67$), that merely sharing rubric structure restores 62\% of total agreement, and that high-quality outputs paradoxically receive the \textit{least} consistent evaluations. \textbf{Second}, we demonstrate that dynamically generating evaluation rubrics grounded in domain knowledge produces more meaningful assessment. We introduce MERG (Metacognitive Enhanced Rubric Generation), a knowledge-driven rubric generation framework whose domain-selective effects confirm this. Agreement \textit{increases} in codified domains (Education +22\%, Academic +27\%) where knowledge anchors evaluators on shared standards, while it decreases in subjective domains where genuine evaluative pluralism emerges. These findings suggest that evaluation rubrics should be dynamically enriched with expert knowledge rather than relying on generic criteria, with implications for reward modeling in RLAIF.

大模型评估评分一致性知识驱动RLAIF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。