arXiv:2609.02942cs.CLcs.AI2026-09

发现评分标准本身能泄露评分信息,影响AI文本评估可靠性

Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

论文配图:Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation
图 1 · 摘自论文原文
  • 仅用评分标准文本训练分类器即可预测评分结果
  • 模型在反事实修改下评分稳定性差,常不更新判断
  • 提示评估体系存在隐藏偏见,适合关注评测可信度的研究者

基于大模型的自动文本评估(LLM-as-a-Judge)广泛依赖评分标准生成判断。我们发现,仅凭评分标准文本,不接触待评文本,分类器就能获得非平凡的预测性能,说明标准本身已隐含可被提取的评价信号。进一步实验表明,当候选文本或评分标准发生反事实变化时,评分模型往往无法可靠调整其判断。这些结果质疑了当前基于标准的自动化评估的可靠性,呼吁对大模型评测方法进行更深入的方法学研究。

原文摘要 · Abstract (English)

LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only on rubric text, without access to any evaluated response, achieve nontrivial predictive performance on judge outputs. This suggests that rubric formulations encode recoverable evaluative signals, allowing scores to be partially anticipated independently of model outputs. Finally, counterfactual perturbations reveal that judges often fail to reliably update their decisions when either the candidate response or the rubric criterion is reversed. Our findings raise concerns about the reliability of rubric-based LLM evaluation and highlight the need for further methodological study of automated evaluation via LLMs.

大模型评测评估偏差自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。