arXiv:2502.11689cs.CL2025-02EMNLP被引 46

用少数据训练大模型做裁判,效果更好更通用。

Improve LLM-as-a-Judge Ability as a General Ability

  • 分两阶段训练:先微调再偏好优化,提升判断力
  • 仅需2%-40%数据量,就达到当前最佳表现
  • 适合想高效训练裁判模型的研究者和开发者

LLM-as-a-Judge 利用大语言模型的生成与推理能力,在多种场景下评估模型输出,提供准确的偏好信号,对对齐人类价值观、确保符合社会规范的可靠AI输出至关重要。现有方法多依赖大量数据或准确性不足,且仅聚焦裁判能力。本文将裁判能力视为LLM的通用能力,提出两阶段训练方法:监督微调(SFT)预热 + 直接偏好优化(DPO)增强,实现风格适配与判断精度提升。同时引入高效数据合成方法生成判别内容。实验表明,本方法仅需其他方法2%至40%的数据量,就在RewardBench上取得最先进性能。此外,该训练方式通过构建复杂裁判任务,增强了模型整体能力,其提供的裁判信号显著提升了内部模型在下游DPO训练中的表现。模型权重与训练数据已开源,以促进后续研究。

原文摘要 · Abstract (English)

LLM-as-a-Judge leverages the generative and reasoning capabilities of large language models (LLMs) to evaluate LLM responses across diverse scenarios, providing accurate preference signals. This approach plays a vital role in aligning LLMs with human values, ensuring ethical and reliable AI outputs that align with societal norms. Recent studies have raised many methods to train LLM as generative judges, but most of them are data consuming or lack accuracy, and only focus on LLM's judge ability. In this work, we regard judge ability as a general ability of LLM and implement a two-stage training approach, comprising supervised fine-tuning (SFT) warm-up and direct preference optimization (DPO) enhancement, to achieve judge style adaptation and improve judgment accuracy. Additionally, we introduce an efficient data synthesis method to generate judgmental content. Experimental results demonstrate that our approach, utilizing only about 2% to 40% of the data required by other methods, achieves SOTA performance on RewardBench. Furthermore, our training method enhances the general capabilities of the model by constructing complicated judge task, and the judge signals provided by our model have significantly enhanced the downstream DPO training performance of our internal models in our test to optimize policy model with Judge Model. We also open-source our model weights and training data to facilitate further research.

大模型评估偏好优化高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。