用大模型当评分员,自动评估文本质量,省时又高效。
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

- 让大模型根据自然语言输出打分,替代人工评测
- 涵盖功能、方法、应用、评估与局限五大方面
- 适合研究者和工程师快速了解该领域全貌
大语言模型的快速发展推动其在多个领域的广泛应用。其中最具前景的应用之一是作为基于自然语言响应的评估者,即“大模型当裁判”(LLMs-as-judges)。该范式因其出色的评估效果、跨任务泛化能力以及以自然语言形式呈现的可解释性,受到学术界与工业界的广泛关注。本文从五个关键角度对这一范式进行全面综述:功能、方法、应用、元评估与局限性。首先系统定义了大模型当裁判的概念并阐述其用途;接着探讨如何构建基于大模型的评估体系;进一步分析其潜在应用场景;讨论在不同情境下评估大模型裁判的方法;最后深入剖析其局限性并展望未来方向。通过结构化、全面的分析,旨在为研究与实践提供洞见。相关资源将持续更新于 https://github.com/CSHaitao/Awesome-LLMs-as-Judges。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has driven their expanding application across various fields. One of the most promising applications is their role as evaluators based on natural language responses, referred to as ''LLMs-as-judges''. This framework has attracted growing attention from both academia and industry due to their excellent effectiveness, ability to generalize across tasks, and interpretability in the form of natural language. This paper presents a comprehensive survey of the LLMs-as-judges paradigm from five key perspectives: Functionality, Methodology, Applications, Meta-evaluation, and Limitations. We begin by providing a systematic definition of LLMs-as-Judges and introduce their functionality (Why use LLM judges?). Then we address methodology to construct an evaluation system with LLMs (How to use LLM judges?). Additionally, we investigate the potential domains for their application (Where to use LLM judges?) and discuss methods for evaluating them in various contexts (How to evaluate LLM judges?). Finally, we provide a detailed analysis of the limitations of LLM judges and discuss potential future directions. Through a structured and comprehensive analysis, we aim aims to provide insights on the development and application of LLMs-as-judges in both research and practice. We will continue to maintain the relevant resource list at https://github.com/CSHaitao/Awesome-LLMs-as-Judges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。