arXiv:2601.21816cs.LG2026-01中稿 · ICML被引 5

基于偏好数据非参数评估大模型性能,更准更可靠。

Nonparametric LLM Evaluation from Preference Data

  • 用去偏机器学习构建非参数框架,不依赖强假设。
  • 支持复杂人类反馈(如平局),并提供可信的不确定性度量。
  • 适合需要精准排名的评测场景,尤其适配预算有限的数据收集。

从人类偏好数据评估大语言模型性能对生成模型排行榜至关重要。现有方法或依赖严格参数假设,或在使用灵活机器学习方法时缺乏有效的不确定性量化。本文提出一种名为DMLRank的非参数统计框架,利用去偏机器学习(DML)比较和排序大模型。引入广义平均排名得分(GARS),可推广常见排名模型(如Bradley-Terry、PageRank/秩中心性),并处理包含平局等复杂人类反馈。DMLRank具有四大优势:(i) 产生高效的GARS估计;(ii) 可自然融合黑箱机器学习方法;(iii) 可结合预训练大模型评估器(如LLM-as-a-judge);(iv) 在预算约束下建议最优偏好数据采集策略。通过合成与真实世界偏好数据集的理论与实证验证,证明了该框架的有效性。总体而言,本框架为模型评测提供了强大、前沿的比较与排序工具。

原文摘要 · Abstract (English)

Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards. However, many existing approaches either rely on restrictive parametric assumptions or lack valid uncertainty quantification when flexible machine learning methods are used. In this paper, we propose a nonparametric statistical framework, called DMLRank, for comparing and ranking LLMs from preference data using debiased machine learning (DML). For this, we introduce generalized average ranking scores (GARS), which generalize commonly used ranking models, including the Bradley-Terry model or PageRank/ Rank centrality, with complex human responses such as ties. DMLRank comes with the following advantages: (i)~It produces statistically efficient estimates of GARS ranking scores. (ii) It naturally allows the incorporation of black-box machine learning methods for estimation. (iii) It can be combined with pre-trained LLM evaluators (e.g., using LLM-as-a-judge). (iv) It suggests optimal policies for collecting preference data under budget constraints. We demonstrate these advantages both theoretically and empirically using both synthetic and real-world preference datasets. In summary, our framework provides practitioners with powerful, state-of-the-art methods for comparing or ranking LLMs for leaderboards.

大模型评估偏好学习非参数统计排名算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。