EasyJudge让普通人也能轻松评估大模型输出质量
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
- 用优化的提示和数据集训练轻量评估模型
- 在消费级显卡上运行,评估结果接近人工与闭源模型
- 提供可视化界面,适合资源有限的研究者使用
近期越来越多研究采用大语言模型(LLMs)来评估其他大模型的输出质量。许多研究依赖闭源模型,主要使用GPT-4作为评估器,但其封闭性导致透明度低、可控性差且成本高。部分研究转向微调开源大模型作为评估器,但现有开源评估模型普遍缺乏友好的可视化工具,也未针对加速推理进行优化,给资源有限或跨领域研究者带来不便。本文提出EasyJudge,一种轻量、精准、高效且用户友好的大模型响应评估工具。它通过精细化数据集和优化提示实现模型优化,在保持与人类及专有模型评估高度一致的同时,支持在消费级GPU甚至CPU上高效运行。我们还提供了详尽分析与案例研究,进一步揭示该方法潜力。
原文摘要 · Abstract (English)
Recently, there has been a growing trend of employing large language models (LLMs) to judge the quality of other LLMs. Many studies have adopted closed-source models, mainly using GPT-4 as the evaluator. However, due to the closed-source nature of the GPT-4 model, employing it as an evaluator has resulted in issues including transparency, controllability, and cost-effectiveness. Some researchers have turned to using fine-tuned open-source LLMs as evaluators. However, existing open-source evaluation LLMs generally lack a user-friendly visualization tool, and they have not been optimized for accelerated model inference, which causes inconvenience for researchers with limited resources and those working across different fields. This paper presents EasyJudge, a model developed to evaluate significant language model responses. It is lightweight, precise, efficient, and user-friendly, featuring an intuitive visualization interface for ease of deployment and use. EasyJudge uses detailed datasets and refined prompts for model optimization, achieving strong consistency with human and proprietary model evaluations. The model optimized with quantitative methods enables EasyJudge to run efficiently on consumer-grade GPUs or even CPUs. We also provide detailed analysis and case studies to further reveal the potential of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。