arXiv:2411.14483cs.CLcs.AI2024-11ACL被引 8

比较大模型输出优劣,用比赛打分方式科学排序。

Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat

  • 通过人类两两对比输出,用爱洛算法构建模型排名
  • 发现评分稳定性受样本数量和评判标准影响显著
  • 提供选型指南,适合资源有限或追求精度的评估场景

选择合适的大型语言模型(LLM)是一项复杂挑战。成对排名已成为评估人类偏好于LLM输出的新方法:由人类根据预设标准对两个模型输出进行两两比较。通过收集这些对比结果,可使用如爱洛(Elo)等算法构建排名。然而,直接将此类算法应用于LLM评估会带来若干挑战。本文系统研究了头对头比较中排名系统的有效性,明确定义了一套基础原则以实现有效排名,并在多个场景下评估了多种排名算法的鲁棒性。分析揭示了影响排名准确性和效率的关键因素,为不同评估情境与资源约束下的方法选择提供了实用指导。

原文摘要 · Abstract (English)

Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a predefined criterion. By collecting these comparisons, a ranking can be constructed using methods such as Elo. However, applying these algorithms as constructed in the context of LLM evaluation introduces several challenges. In this paper, we explore the effectiveness of ranking systems for head-to-head comparisons of LLMs. We formally define a set of fundamental principles for effective ranking and conduct a series of extensive evaluations on the robustness of several ranking algorithms in the context of LLMs. Our analysis uncovers key insights into the factors that affect ranking accuracy and efficiency, offering guidelines for selecting the most appropriate methods based on specific evaluation contexts and resource constraints.

大模型评测排名算法人机评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。