arXiv:2508.02994cs.AI2025-08被引 19

用AI当裁判评估大模型,提升评测效率与深度。

When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs

  • 让AI代理通过推理和多角度分析来评判其他模型输出。
  • 相比人工评测,该方法在成本和可扩展性上优势明显。
  • 适合需要大规模、高复杂度评测的科研与产业场景。

随着大语言模型能力与自主性的提升,对其在开放性与复杂任务中输出的评估已成为关键瓶颈。一种新范式正在兴起:使用AI代理作为评估者本身。‘代理即裁判’方法利用大模型的推理与共情能力,评估其他模型输出的质量与安全性,为人工评估提供了可扩展且细致的替代方案。本文定义了该概念,梳理其从单模型裁判到动态多代理辩论框架的发展历程,系统比较了不同方法在可靠性、成本与人类对齐方面的表现,并调研了其在医疗、法律、金融及教育等领域的实际应用。最后,我们指出当前面临的挑战,包括偏见、鲁棒性与元评估问题,并提出未来研究方向。综上,基于代理的评估可补充但不可替代人类监督,是迈向下一代大模型可信、可扩展评测的重要一步。

原文摘要 · Abstract (English)

As large language models (LLMs) grow in capability and autonomy, evaluating their outputs-especially in open-ended and complex tasks-has become a critical bottleneck. A new paradigm is emerging: using AI agents as the evaluators themselves. This "agent-as-a-judge" approach leverages the reasoning and perspective-taking abilities of LLMs to assess the quality and safety of other models, promising calable and nuanced alternatives to human evaluation. In this review, we define the agent-as-a-judge concept, trace its evolution from single-model judges to dynamic multi-agent debate frameworks, and critically examine their strengths and shortcomings. We compare these approaches across reliability, cost, and human alignment, and survey real-world deployments in domains such as medicine, law, finance, and education. Finally, we highlight pressing challenges-including bias, robustness, and meta evaluation-and outline future research directions. By bringing together these strands, our review demonstrates how agent-based judging can complement (but not replace) human oversight, marking a step toward trustworthy, scalable evaluation for next-generation LLMs.

AI评估大模型多代理评测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。