用大模型当裁判评估其他大模型质量,效果接近人工且更高效。
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
- 用不同大模型+定制提示词充当裁判,自动评分其他模型输出。
- GPT-4o和32B以上开源模型在多数任务中与人工判断相关性超0.8。
- 小模型如Qwen2.5-14B也表现良好,适合资源有限场景使用。
以大语言模型(LLM)作为裁判,评估目标机器学习模型(尤其是大模型)的输出质量。该方法通过特定设计的评判提示词,将一个模型转化为自动化评估工具,相比人工评审更快速、一致。本研究测试了37个不同规模的对话型大模型,搭配5种提示词,引入二级裁判机制,并使用5个微调后的模型进行评估。针对8类评估任务构建了基于人工标注的基准数据集。实证结果表明,在采用合适提示词时,大模型裁判与人类判断高度相关,尤其在GPT-4o、参数量≥32B的开源模型及部分小型模型(如Qwen2.5-14B)上表现优异。
原文摘要 · Abstract (English)
A Large Language Model (LLM) as judge evaluates the quality of victim Machine Learning (ML) models, specifically LLMs, by analyzing their outputs. An LLM as judge is the combination of one model and one specifically engineered judge prompt that contains the criteria for the analysis. The resulting automation of the analysis scales up the complex evaluation of the victim models' free-form text outputs by faster and more consistent judgments compared to human reviewers. Thus, quality and security assessments of LLMs can cover a wide range of the victim models' use cases. Being a comparably new technique, LLMs as judges lack a thorough investigation for their reliability and agreement to human judgment. Our work evaluates the applicability of LLMs as automated quality assessors of victim LLMs. We test the efficacy of 37 differently sized conversational LLMs in combination with 5 different judge prompts, the concept of a second-level judge, and 5 models fine-tuned for the task as assessors. As assessment objective, we curate datasets for eight different categories of judgment tasks and the corresponding ground-truth labels based on human assessments. Our empirical results show a high correlation of LLMs as judges with human assessments, when combined with a suitable prompt, in particular for GPT-4o, several open-source models with $\geqslant$ 32B parameters, and a few smaller models like Qwen2.5 14B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。