用大模型当裁判评估AI,解决传统方法难应对开放场景的问题
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- 用大模型自动打分、排序或筛选任务结果
- 提出三维度分类框架:评什么、怎么评、如何验证
- 适合研究评估机制或想高效测试AI系统的学者
评估与评测一直是人工智能和自然语言处理中的关键挑战。传统方法多基于匹配或小型模型,在开放性和动态性场景下表现不足。近年来,大语言模型的发展催生了‘大模型作为裁判’(LLM-as-a-judge)的新范式,即利用大模型对各类机器学习任务的结果进行评分、排序或选择。本文对基于大模型的判断与评估进行全面综述,从输入与输出双视角定义该范式;构建涵盖‘评判对象’、‘评判方式’和‘基准方法’三个维度的系统性分类体系;并指出当前核心挑战与未来发展方向。更多资源可访问:https://llm-as-a-judge.github.io 及 https://github.com/llm-as-a-judge/Awesome-LLM-as-a-judge。
原文摘要 · Abstract (English)
Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios. Recent advancements in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm, where LLMs are leveraged to perform scoring, ranking, or selection for various machine learning evaluation scenarios. This paper presents a comprehensive survey of LLM-based judgment and assessment, offering an in-depth overview to review this evolving field. We first provide the definition from both input and output perspectives. Then we introduce a systematic taxonomy to explore LLM-as-a-judge along three dimensions: what to judge, how to judge, and how to benchmark. Finally, we also highlight key challenges and promising future directions for this emerging area. More resources on LLM-as-a-judge are on the website: https://llm-as-a-judge.github.io and https://github.com/llm-as-a-judge/Awesome-LLM-as-a-judge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。