arXiv:2602.08672cs.CLcs.LG2026-02Conference of the …被引 16

让大模型自己设计评分标准,评估生成文本质量。

Learning to Judge: LLMs Designing and Applying Evaluation Rubrics

  • 大模型自主生成可解释的评价维度,任务相关性强。
  • 在事实类任务中评分一致性下降,开放模型表现弱于闭源模型。
  • 适合关注模型评估机制与人机评价对齐的研究者。

大语言模型(LLMs)越来越多地被用作自然语言生成的评估工具,采用人类定义的评分标准来评估系统输出。然而,人类制定的标准往往静态且与模型内部的语言质量表征不一致。本文提出GER-Eval(生成评估标准用于评估),探究大模型能否自主设计并应用自己的评价标准。我们评估了模型自定义标准的语义连贯性、评分可靠性及其与人类标准的一致性。结果显示,大模型能可靠生成可解释且任务感知的评价维度,并在模型内部保持一致应用,但在涉及事实和知识的任务中评分可靠性下降。闭源模型如GPT-4o在评分一致性与跨模型泛化能力上优于开放权重模型如Llama。研究揭示评估是大模型的一种可学习语言能力,模型内一致但模型间割裂,呼吁发展联合建模人类与大模型评价语言的新方法,以提升可靠性与可解释性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as evaluators for natural language generation, applying human-defined rubrics to assess system outputs. However, human rubrics are often static and misaligned with how models internally represent language quality. We introduce GER-Eval (Generating Evaluation Rubrics for Evaluation) to investigate whether LLMs can design and apply their own evaluation rubrics. We evaluate the semantic coherence and scoring reliability of LLM-defined criteria and their alignment with human criteria. LLMs reliably generate interpretable and task-aware evaluation dimensions and apply them consistently within models, but their scoring reliability degrades in factual and knowledge-intensive settings. Closed-source models such as GPT-4o achieve higher agreement and cross-model generalization than open-weight models such as Llama. Our findings position evaluation as a learned linguistic capability of LLMs, consistent within models but fragmented across them, and call for new methods that jointly model human and LLM evaluative language to improve reliability and interpretability.

大模型评估评分标准语言质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。