首个开源全能型模型评估器,支持评分、对比与批评生成。
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
- 集成评分、对比、格式化评估和批评生成的统一模型
- 在新基准JudgerBench上验证,覆盖多种主观评估任务
- 适合研究者快速搭建自动化评估系统
高效准确的评估对大语言模型的持续优化至关重要。尽管主观评估更贴近真实使用场景和人类偏好,但人工评估成本高且不可复现,因此精准的自动化评估器(评判模型)尤为关键。本文提出首个开源的全功能评判模型CompassJudger-1,它是一个通用大模型,具备四项能力:1)作为奖励模型进行单个评分和双模型对比;2)按指定格式执行评估;3)生成批评意见;4)完成多样任务如通用语言模型。为统一评估不同评判模型的表现,我们还构建了新基准JudgerBench,涵盖多种主观评估任务及广泛主题。CompassJudger-1 提供全面评估解决方案,同时保持灵活性以适应多样化需求。相关工具已在 https://github.com/open-compass/CompassJudger 开源,期待推动评估方法研究合作与进展。
原文摘要 · Abstract (English)
Efficient and accurate evaluation is crucial for the continuous improvement of large language models (LLMs). Among various assessment methods, subjective evaluation has garnered significant attention due to its superior alignment with real-world usage scenarios and human preferences. However, human-based evaluations are costly and lack reproducibility, making precise automated evaluators (judgers) vital in this process. In this report, we introduce \textbf{CompassJudger-1}, the first open-source \textbf{all-in-one} judge LLM. CompassJudger-1 is a general-purpose LLM that demonstrates remarkable versatility. It is capable of: 1. Performing unitary scoring and two-model comparisons as a reward model; 2. Conducting evaluations according to specified formats; 3. Generating critiques; 4. Executing diverse tasks like a general LLM. To assess the evaluation capabilities of different judge models under a unified setting, we have also established \textbf{JudgerBench}, a new benchmark that encompasses various subjective evaluation tasks and covers a wide range of topics. CompassJudger-1 offers a comprehensive solution for various evaluation tasks while maintaining the flexibility to adapt to diverse requirements. Both CompassJudger and JudgerBench are released and available to the research community athttps://github.com/open-compass/CompassJudger. We believe that by open-sourcing these tools, we can foster collaboration and accelerate progress in LLM evaluation methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。