选错基准模型会严重影响大模型评估结果,中等水平的模型才是好基准。
Mediocrity is the key for LLM as a Judge Anchor Selection
- 用22个不同模型做基准,发现极端表现者不适合作为评估锚点
- 标准评测规模不足,难以区分性能相近模型
- 推荐选用中等水平模型作基准,提升评估可靠性
LLM作为评判者已成为开放生成任务评估的标准方法。为应对成对比较带来的二次复杂度问题,主流基准如Arena-Hard和AlpacaEval通常将所有模型与单一锚模型进行比较。然而,锚模型选择对评估结果可靠性的影响尚未被充分研究。本文在Arena-Hard-v2.0数据集上系统评估了22种不同锚模型,发现锚选择至关重要:劣质锚模型会显著降低与人类评分的相关性。我们发现,常见选择(表现最好或最差的模型)效果不佳,因其始终优于或劣于其他模型,难以反映相对排名。进一步量化显示,锚选择的影响程度与裁判模型选择相当。据此提出两项建议:一是进行功效分析,计算锚基评估所需的足够样本量,发现当前标准规模不足以可靠区分性能接近的模型;二是提供选择有效锚模型的指南,确保评估的可靠性与效率。
原文摘要 · Abstract (English)
The ``LLM-as-a-judge'' paradigm has become a standard method for evaluating open-ended generation. To address the quadratic scalability costs of pairwise comparisons, popular benchmarks like Arena-Hard and AlpacaEval compare all models against a single anchor. However, despite its widespread use, the impact of anchor selection on the reliability of the results remains largely unexplored. In this work, we systematically investigate the effect of anchor selection by evaluating 22 different anchors on the Arena-Hard-v2.0 dataset. We find that the choice of anchor is critical: a poor anchor can dramatically reduce correlation with human rankings. We identify that common anchor choices (best-performing and worst-performing models) make poor anchors. Because these extreme anchors are consistently better or worse than all other models, they are seldom indicative of the relative ranking of the models. We further quantify the effect size of anchor selection, showing it is comparable to the selection of a judge model. We conclude with actionable recommendations. First, we conduct a power analysis, and compute sufficient benchmark sizes for anchor-based evaluation, finding that standard benchmark sizes are insufficient for pairwise evaluation and fail to distinguish between competitive models reliably. Second, we provide guidelines for selecting informative anchors to ensure reliable and efficient evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。