训练一个能精准评阅的AI裁判模型,让大模型评估更可靠。
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
- 设计情境化提示与指令生成方法,提升评估灵活性
- 在两个人工标注基准上实现与人类偏好高度一致的评分
- 提供数据平衡、多目标训练等实用指南,适合评估研究者
大型语言模型(LLM)的快速发展为将其作为评价裁判提供了新可能。本文介绍Themis——一个经过微调的LLM裁判模型,能够实现上下文感知的复杂评估。我们系统梳理了Themis的开发流程,提出基于场景的评估提示和两种新颖的受控指令生成方法,使模型能有效从教师模型中蒸馏评估能力,同时保持持续优化的灵活性。我们构建了两个用于元评估的人工标注基准,证明Themis可在经济成本下实现与人类偏好高度对齐。此外,我们揭示了LLM作为裁判范式中的关键洞察:单纯从强模型进行知识蒸馏,并不保证性能随规模提升;我们提出基于指令遵循难度的缓解策略。最后,我们给出涵盖数据平衡、提示定制、多目标训练和指标聚合的实践建议。我们的方法、发现及配套数据、模型检查点旨在推动该领域的后续研究。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge that delivers sophisticated context-aware evaluations. We provide a comprehensive overview of the development pipeline for Themis, highlighting its scenario-dependent evaluation prompts and two novel methods for controlled instruction generation. These designs enable Themis to effectively distill evaluative skills from teacher models, while retaining flexibility for continuous development. We introduce two human-labeled benchmarks for meta-evaluation, demonstrating that Themis can achieve high alignment with human preferences in an economical manner. Additionally, we explore insights into the LLM-as-a-judge paradigm, revealing nuances in performance and the varied effects of reference answers. Notably, we observe that pure knowledge distillation from strong LLMs, though common, does not guarantee performance improvement through scaling. We propose a mitigation strategy based on instruction-following difficulty. Furthermore, we provide practical guidelines covering data balancing, prompt customization, multi-objective training, and metric aggregation. We aim for our method and findings, along with the fine-tuning data, benchmarks, and model checkpoints, to support future research and development in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。