arXiv:2505.16003cs.CLcs.AI2025-05

用熵最大化方法提升大模型评价与人类判断的一致性。

SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models

  • 基于少量人工偏好数据,通过熵最大化校准评分
  • 在真实场景中实现0.57的斯皮尔曼相关性,优于现有方法
  • 效率提升5-30倍,适合生产级部署

LLM-as-a-Judge范式提供了一种可扩展、无需参考的模型评估方式。尽管已有多种校准技术试图使这类评估器更贴近人类判断,但以往研究主要聚焦于结构化程度高的小规模基准测试。因此,这些校准方法在真实世界开放任务中的泛化能力尚不明确。本文发现,当前最先进校准评估器在开放任务中常表现不佳,与人类判断的相关性弱甚至为负。为此,我们提出SLMEval,一种基于熵最大化的新型高效校准方法,仅需少量人工偏好数据即可估计模型质量的潜在分布,并据此重加权评估分数。SLMEval在两个真实生产场景和公开基准上均表现出与人类判断强相关性。例如,在一项任务中,其斯皮尔曼相关性达0.57,而G-Eval为负相关。此外,相比GPT-4驱动的校准方法如G-eval,SLMEval将评估成本降低了5-30倍。

原文摘要 · Abstract (English)

The LLM-as-a-Judge paradigm offers a scalable, reference-free approach for evaluating language models. Although several calibration techniques have been proposed to better align these evaluators with human judgment, prior studies focus primarily on narrow, well-structured benchmarks. As a result, it remains unclear whether such calibrations generalize to real-world, open-ended tasks. In this work, we show that SOTA calibrated evaluators often fail in these settings, exhibiting weak or even negative correlation with human judgments. To address this, we propose SLMEval, a novel and efficient calibration method based on entropy maximization over a small amount of human preference data. By estimating a latent distribution over model quality and reweighting evaluator scores accordingly, SLMEval achieves strong correlation with human evaluations across two real-world production use cases and the public benchmark. For example, on one such task, SLMEval achieves a Spearman correlation of 0.57 with human judgments, while G-Eval yields a negative correlation. In addition, SLMEval reduces evaluation costs by 5-30x compared to GPT-4-based calibrated evaluators such as G-eval.

大模型评估熵校准人类对齐低成本评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。