arXiv:2412.05506stat.MLcs.LG2024-12被引 3

用置信图评估大模型排名的不确定性,更可靠地比较医学领域表现。

Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation

  • 基于非参数方法构建置信图,捕捉模型排名的不确定性。
  • 在真实医疗数据上验证,可区分不同模型在多个领域的优劣。
  • 适合关注大模型评估可靠性、需稳健排序的研究者使用。

我们研究大语言模型(LLMs)的排名推断问题。对齐是缓解大模型幻觉的关键挑战,基于 best-of-$N$ 策略的模型排名已被证明能有效提升对齐性。本文提出一种新的假设检验框架,用于语言模型排名之间的统计推断。该框架基于非参数上下文排名机制,用于评估模型在特定领域的能力,利用非参数评分方法以反映其对提示的敏感性。为刻画排名的组合复杂性,我们引入新颖的置信图概念,通过哈斯图(Hasse diagram)将所有可能的排名置信集统一表示为一个有向图。我们通过扩展高斯乘子自助法理论,证明了所提置信图的有效性,该理论适用于独立但不等分布的样本最大经验过程。在合成数据和真实医疗数据上的大量实验表明,该方法为不同大模型在多个医学领域中的性能评估提供了深刻洞见。

原文摘要 · Abstract (English)

We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs has proven to be an effective tool to improve alignment based on the best-of-$N$ policy. In this paper, we propose a new inferential framework for hypothesis testing among the ranking for language models. Our framework is based on a nonparametric contextual ranking framework designed to assess large language models' domain-specific expertise, leveraging nonparametric scoring methods to account for their sensitivity to the prompts. To characterize the combinatorial complexity of the ranking, we introduce a novel concept of confidence diagram, which leverages a Hasse diagram to represent the entire confidence set of rankings by a single directed graph. We show the validity of the proposed confidence diagram by advancing the Gaussian multiplier bootstrap theory to accommodate the supremum of independent empirical processes that are not necessarily identically distributed. Extensive numerical experiments conducted on both synthetic and real data demonstrate that our approach offers valuable insight into the evaluation for the performance of different LLMs across various medical domains.

大模型评估置信图非参数统计医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。