arXiv:2601.17312cs.CLcs.AI2026-01被引 3

用大模型评估大模型,解决评估不靠谱的问题。

Meta-Judging with Large Language Models: Concepts, Methods, and Challenges

  • 让大模型当评委的评委,提升评估稳定性。
  • 发现原评估方式存在提示敏感、偏见等缺陷。
  • 适合研究自动化评测与模型对齐的学者。

大语言模型(LLMs)发展迅速,现常被用作评估者,即所谓的‘大模型作为裁判’(LLM-as-a-Judge),用于对模型输出质量进行评估。然而,近期研究表明此类评估存在显著缺陷,包括对提示词敏感、系统性偏见、冗长性影响及不可靠或虚构的推理过程。这些局限促使了更稳健范式——‘大模型作为元裁判’(LLM-as-a-Meta-Judge)的发展。本文综述了元裁判的最新进展,提出了涵盖六个关键视角的框架:(i) 概念基础,(ii) 元裁判机制,(iii) 对齐训练方法,(iv) 评估方法,(v) 局限性与失效模式,(vi) 未来方向。通过分析LLM-as-a-Judge的不足并总结元裁判的最新成果,我们认为该范式为更稳定、可信的自动化评估提供了前景,同时指出了成本、提示敏感性和共享模型偏见等仍需解决的挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) are evolving fast and are now frequently used as evaluators, in a process typically referred to as LLM-as-a-Judge, which provides quality assessments of model outputs. However, recent research points out significant vulnerabilities in such evaluation, including sensitivity to prompts, systematic biases, verbosity effects, and unreliable or hallucinated rationales. These limitations motivated the development of a more robust paradigm, dubbed LLM-as-a-Meta-Judge. This survey reviews recent advances in meta-judging and organizes the literature, by introducing a framework along six key perspectives: (i) Conceptual Foundations, (ii) Mechanisms of Meta-Judging, (iii) Alignment Training Methods, (iv) Evaluation, (v) Limitations and Failure Modes, and (vi) Future Directions. By analyzing the limitations of LLM-as-a-Judge and summarizing recent advances in meta-judging by LLMs, we argue that LLM-as-a-Meta-Judge offers a promising direction for more stable and trustworthy automated evaluation, while highlighting remaining challenges related to cost, prompt sensitivity, and shared model biases, which must be addressed to advance the next generation of LLM evaluation methodologies.

大模型评估元裁判自动化评测可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。