arXiv:2601.05111cs.CLcs.AI2026-01被引 17

用智能体替代大模型做评价,让评估更可靠、可验证。

Agent-as-a-Judge

  • 引入智能体机制,支持规划与工具验证
  • 可多轮推理并协作,提升评估深度
  • 适合复杂任务评估,如科研与工程

LLM-as-a-Judge 通过大语言模型实现大规模评估,但面对日益复杂的多步任务时,其可靠性受限于固有偏见、浅层单次推理及缺乏真实世界验证能力。为此,研究转向 Agent-as-a-Judge 模式,利用规划、工具增强验证、多智能体协作与持续记忆,实现更鲁棒、可验证、细致的评估。尽管智能体评估系统快速涌现,该领域仍缺乏统一框架。本文首次系统性综述这一演进,识别关键维度,建立发展分类体系,梳理通用与专业领域的核心方法与应用,并分析前沿挑战与未来方向,为下一代智能体评估提供清晰路线图。

原文摘要 · Abstract (English)

LLM-as-a-Judge has revolutionized AI evaluation by leveraging large language models for scalable assessments. However, as evaluands become increasingly complex, specialized, and multi-step, the reliability of LLM-as-a-Judge has become constrained by inherent biases, shallow single-pass reasoning, and the inability to verify assessments against real-world observations. This has catalyzed the transition to Agent-as-a-Judge, where agentic judges employ planning, tool-augmented verification, multi-agent collaboration, and persistent memory to enable more robust, verifiable, and nuanced evaluations. Despite the rapid proliferation of agentic evaluation systems, the field lacks a unified framework to navigate this shifting landscape. To bridge this gap, we present the first comprehensive survey tracing this evolution. Specifically, we identify key dimensions that characterize this paradigm shift and establish a developmental taxonomy. We organize core methodologies and survey applications across general and professional domains. Furthermore, we analyze frontier challenges and identify promising research directions, ultimately providing a clear roadmap for the next generation of agentic evaluation.

智能体评估大模型评测多智能体评价系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。