arXiv:2504.17087cs.AI2025-04被引 22

用大模型当裁判的裁判,提升AI评估结果的可靠性。

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

  • 设计多智能体框架,让多个大模型协作评判其他模型输出。
  • 在JudgeBench数据集上比原始判断提升15.55%,比单模型基准高8.37%。
  • 适合需要高质量评估的AI训练与评测场景,如强化学习对齐。

大语言模型在各领域广泛应用,但任务复杂度上升使得其输出评估愈发困难。相较于人类评价者,使用大模型辅助评估更具效率。然而,现有研究多聚焦于对齐模型判断与人类偏好,忽视了人类判断中存在的偏见与错误。此外,如何从多个可能的模型判断中选取合适结果仍不明确。为此,本文提出一个三阶段元裁判选择流程:1)结合GPT-4与人工专家制定全面评分标准;2)由三个先进大模型智能体进行打分;3)设定阈值剔除低分判断。相比以单一模型同时担任裁判与元裁判的方法,本方案引入多智能体协作与更全面的评分体系。在JudgeBench数据集上的实验表明,该方法相较原始判断提升约15.55%,相较单智能体基线提升约8.37%。本工作展示了大模型作为元裁判的潜力,为构建用于大模型作为裁判的强化学习偏好数据集奠定了基础。

原文摘要 · Abstract (English)

Large language models (LLMs) are being widely applied across various fields, but as tasks become more complex, evaluating their responses is increasingly challenging. Compared to human evaluators, the use of LLMs to support performance evaluation offers a more efficient alternative. However, most studies focus mainly on aligning LLMs' judgments with human preferences, overlooking the existence of biases and mistakes in human judgment. Furthermore, how to select suitable LLM judgments given multiple potential LLM responses remains underexplored. To address these two aforementioned issues, we propose a three-stage meta-judge selection pipeline: 1) developing a comprehensive rubric with GPT-4 and human experts, 2) using three advanced LLM agents to score judgments, and 3) applying a threshold to filter out low-scoring judgments. Compared to methods using a single LLM as both judge and meta-judge, our pipeline introduces multi-agent collaboration and a more comprehensive rubric. Experimental results on the JudgeBench dataset show about 15.55\% improvement compared to raw judgments and about 8.37\% improvement over the single-agent baseline. Our work demonstrates the potential of LLMs as meta-judges and lays the foundation for future research on constructing preference datasets for LLM-as-a-judge reinforcement learning.

大模型评估多智能体元裁判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。