arXiv:2601.08654cs.CLcs.AI2026-01被引 9

用可追溯证据和校准机制提升大模型评分可靠性

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

  • 将人类评分标准转化为固定任务规范,确保评分一致
  • 在多个基准上实现比现有方法更高的真人评分一致性
  • 适合需要可信、可审计文本评估的场景

基于评分量表的文本评估越来越多地使用大语言模型作为可扩展的评判者,但将冻结的黑箱模型与人类评分标准对齐仍具挑战。我们将其视为一种准则迁移问题:目标不仅是让LLM打分,而是将人类评分意图转化为稳定、可审计且与人类对齐的评分协议。我们识别出三类常见失败模式:评分执行漂移、不可验证的分数归因、与人类尺度的错位。为此,我们提出Rulers,一种三阶段推理时框架,实现可靠、基于证据的评分量表文本评估。Rulers首先将人类评分量表转为锁定的任务级规范,再通过结构化检查清单决策、类型化证据支持及可选的摘录验证执行规范,最后进行后验校准以对齐模型信号与人类分数边界。在涵盖论文评分、摘要评估、英语作为第二语言写作评价和结构化输入生成的四个量表驱动基准上,Rulers在多个冻结主干模型下多数情况下达到更强的人类评分一致性。进一步分析表明,Rulers更符合真实人类评分分布,在语义等价的量表扰动下更具稳定性,且三个组件均带来收益。结果表明,可靠的LLM评判需固定准则、可追溯证据和校准的分数解释,而非仅靠提示词设计。代码已公开于https://anonymous.4open.science/r/Rulers_0525-3328。

原文摘要 · Abstract (English)

Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer problem: the goal is not merely to prompt an LLM to assign a score, but to transfer human rubric intent into a stable, auditable, and human-aligned scoring protocol. We identify three recurring failure modes in LLM-based rubric scoring: rubric execution drift, unverifiable score attribution, and human-scale misalignment. To address these failure modes, we introduce Rulers, a three-stage inference-time framework for reliable, evidence-grounded rubric-based text evaluation. Rulers first converts a human rubric into a locked task-level specification, then executes the specification with structured checklist decisions, typed evidence grounding, and extractive quote verification when applicable, and finally applies post-hoc calibration to align model-derived signals with human score boundaries. Across four rubric-governed benchmarks covering essay scoring, summarization assessment, EFL writing evaluation, and structured-input text generation, Rulers achieves stronger human-score agreement in most evaluated settings across multiple frozen backbone models. Further analyses show that Rulers better matches empirical human score distributions, improves stability under semantically equivalent rubric perturbations, and benefits from each of its three components. These results suggest that reliable LLM judging requires fixed criteria, traceable evidence, and calibrated score interpretation rather than prompt phrasing alone. Our code is available at https://anonymous.4open.science/r/Rulers_0525-3328.

文本评估大模型评判可解释性评分对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。