给大模型排名加不确定性分析,避免误判排名差异。
Prompt-Dependent Ranking of Large Language Models with Uncertainty Quantification
- 用上下文感知的伯莱德-泰勒-卢斯模型建模不同提示下的模型表现
- 通过置信区间生成统计有效的排名置信集,识别真正显著的差距
- 适合关注排名可靠性、需做稳健决策的研究者与工程师
基于成对比较的排名在经济与计算系统中至关重要。对于大语言模型(LLMs),排名通常由人类偏好数据构建,并以排行榜形式指导部署决策。然而现有方法依赖点估计,隐含将排名视为固定对象,忽视了估计噪声和上下文相关的性能波动。据此做出的决策可能导致资源错配和福利损失。本文研究在成对人类偏好下提示相关的排名推断,提出具有统计有效不确定性保证的决策安全排名框架。我们采用上下文伯莱德-泰勒-卢斯模型,使每个模型的潜在效用随输入提示变化。不直接估计效用点值,而是对诱导出的排名进行推断,基于成对效用差的联合置信区间构建置信集。该方法产生针对特定提示的边际与联合置信集。实证上,利用大规模人类偏好数据(来自LLM评估),我们发现排名在不同提示特征间存在显著差异,许多看似明显的排名差异并不具有统计显著性。此外,不确定性感知的排名仅在数据支持时返回完全排序,否则返回部分序。
原文摘要 · Abstract (English)
Rankings derived from pairwise comparisons are central to many economic and computational systems. In the context of large language models (LLMs), rankings are typically constructed from human preference data and presented as leaderboards that guide deployment decisions. However, existing approaches rely on point estimates, implicitly treating rankings as fixed objects despite substantial estimation noise and context-dependent performance variation. Acting on such rankings can lead to misallocation and welfare loss when apparent differences are not statistically meaningful. We study prompt-dependent ranking inference under pairwise human preferences and develop a framework for decision-safe rankings with statistically valid uncertainty guarantees. We model preferences using a contextual Bradley-Terry-Luce model in which the latent utility of each model depends on the input prompt. Rather than targeting point estimates of utilities, we directly conduct inference on induced rankings, constructing confidence sets based on simultaneous confidence intervals for pairwise utility differences. This approach yields statistically valid marginal and simultaneous confidence sets for prompt-specific ranks. Our framework connects recent advances in rank inference to contextual preference learning and provides tools for robust ranking-based decision-making. Empirically, using large-scale human preference data from LLM evaluations, we show that rankings vary substantially across prompt characteristics and that many apparent rank differences are not statistically distinguishable. We further demonstrate how uncertainty-aware rankings identify dominance only when supported by the data and otherwise return partial orders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。