用大模型当裁判,如何保证评分可靠?
A Survey on LLM-as-a-Judge
- 让大模型自动评估复杂任务,替代人工打分
- 提出新基准测试模型评分一致性与公平性
- 适合研究评估系统、AI质检的学者和工程师
准确且一致的评估对众多领域的决策至关重要,但受主观性、差异性和规模限制,仍具挑战。大语言模型(LLMs)在多个领域表现卓越,催生了「大模型作为裁判」(LLM-as-a-Judge)的新范式,即利用大模型对复杂任务进行评价。凭借处理多模态数据、提供可扩展、低成本且一致评估的能力,大模型成为传统专家评审的有力替代方案。然而,保障其评估可靠性仍是关键难题,需精心设计与标准化。本文全面综述了大模型作为裁判的技术路径,核心问题为:如何构建可靠的评估系统?我们探讨提升一致性的策略,包括缓解偏差、适应多样化评估场景,并提出评估系统可靠性的方法论,辅以专为该目标设计的新基准。同时,讨论实际应用、挑战与未来方向,为该快速发展的领域提供基础参考。
原文摘要 · Abstract (English)
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains a challenging task due to inherent subjectivity, variability, and scale. Large Language Models (LLMs) have achieved remarkable success across diverse domains, leading to the emergence of "LLM-as-a-Judge," where LLMs are employed as evaluators for complex tasks. With their ability to process diverse data types and provide scalable, cost-effective, and consistent assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-Judge systems remains a significant challenge that requires careful design and standardization. This paper provides a comprehensive survey of LLM-as-a-Judge, addressing the core question: How can reliable LLM-as-a-Judge systems be built? We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse assessment scenarios. Additionally, we propose methodologies for evaluating the reliability of LLM-as-a-Judge systems, supported by a novel benchmark designed for this purpose. To advance the development and real-world deployment of LLM-as-a-Judge systems, we also discussed practical applications, challenges, and future directions. This survey serves as a foundational reference for researchers and practitioners in this rapidly evolving field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。