arXiv:2511.20604cs.CLcs.AI2025-11NeurIPS被引 7

用大模型当裁判评估对齐性,效果不输人工。

On Evaluating LLM Alignment by Evaluating LLMs as Judges

  • 让大模型自己当裁判,评估其对齐人类偏好能力
  • 新基准AlignEval在排名上媲美或超越现有自动评估方法
  • 发现生成与评价能力高度相关,适合评估模型对齐性

大语言模型(LLMs)的对齐性评估是关键环节,需确保模型回应有用、诚实、安全并准确遵循指令。传统评估依赖人工标注或强模型裁判直接评判模型输出;而近年也有研究将大模型本身作为裁判。本文系统分析了不同大模型在生成与评估能力间的一致性(GE-consistency),发现当以强模型偏好基准(LLM preference oracle)评估时,生成与评价能力呈显著正相关。基于此,我们提出一种新范式:不直接评估模型输出,而是通过评估其作为裁判的表现来衡量对齐性。实验表明,所提基准AlignEval在排序大模型时,能有效捕捉人类偏好,表现匹配或优于AlpacaEval和Arena-Hard等主流自动评估基准。本研究揭示了生成与评价能力的内在关联,并提供了一种无需直接评判输出的新对齐评估方法。

原文摘要 · Abstract (English)

Alignment with human preferences is an important evaluation aspect of LLMs, requiring them to be helpful, honest, safe, and to precisely follow human instructions. Evaluating large language models' (LLMs) alignment typically involves directly assessing their open-ended responses, requiring human annotators or strong LLM judges. Conversely, LLMs themselves have also been extensively evaluated as judges for assessing alignment. In this work, we examine the relationship between LLMs' generation and evaluation capabilities in aligning with human preferences. To this end, we first conduct a comprehensive analysis of the generation-evaluation consistency (GE-consistency) among various LLMs, revealing a strong correlation between their generation and evaluation capabilities when evaluated by a strong LLM preference oracle. Utilizing this finding, we propose a benchmarking paradigm that measures LLM alignment with human preferences without directly evaluating their generated outputs, instead assessing LLMs in their role as evaluators. Our evaluation shows that our proposed benchmark, AlignEval, matches or surpasses widely used automatic LLM evaluation benchmarks, such as AlpacaEval and Arena-Hard, in capturing human preferences when ranking LLMs. Our study offers valuable insights into the connection between LLMs' generation and evaluation capabilities, and introduces a benchmark that assesses alignment without directly evaluating model outputs.

大模型对齐评估基准模型自评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。