arXiv:2409.14664cs.CL2024-09EMNLP被引 30

用正负样本优化模型,让大模型自动生成更可靠的评估结果。

Direct Judgement Preference Optimization

  • 通过偏好优化融合正负样本数据提升评判能力。
  • 13个基准中10项表现最优,超越GPT-4o等强基线。
  • 可抵抗位置/长度偏差,适配任意评估协议,输出改进建议。

自动评估对模型质量判断与迭代至关重要。近期研究尝试训练大语言模型(LLMs)作为生成式评判者,以评估并批评其他模型的输出。本文探索从正负样本中学习的偏好优化方法,以提升生成式评判者在多种应用场景下的评估能力。我们采用三种不同方式收集各场景的偏好对,分别从不同角度改进评判模型。在广泛基准上的综合实验表明该方法有效:我们的生成式评判者在13个基准中10项表现最佳,优于GPT-4o及专用评判模型。进一步分析显示,该模型能稳健应对位置偏倚、长度偏倚等固有偏差,灵活适应从业者指定的评估协议,并为下游生成模型提供有价值的语言反馈。

原文摘要 · Abstract (English)

Auto-evaluation is crucial for assessing response quality and offering feedback for model development. Recent studies have explored training large language models (LLMs) as generative judges to evaluate and critique other models' outputs. In this work, we investigate the idea of learning from both positive and negative data with preference optimization to enhance the evaluation capabilities of LLM judges across an array of different use cases. We achieve this by employing three approaches to collect the preference pairs for different use cases, each aimed at improving our generative judge from a different perspective. Our comprehensive study over a wide range of benchmarks demonstrates the effectiveness of our method. In particular, our generative judge achieves the best performance on 10 out of 13 benchmarks, outperforming strong baselines like GPT-4o and specialized judge models. Further analysis show that our judge model robustly counters inherent biases such as position and length bias, flexibly adapts to any evaluation protocol specified by practitioners, and provides helpful language feedback for improving downstream generator models.

大模型评估偏好优化生成式评判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。