arXiv:2410.11594cs.LGcs.AI2024-10被引 17

给大模型评分的不确定性做量化,提升评估可信度

Black-box Uncertainty Quantification Method for LLM-as-a-Judge

  • 通过分析评分与可能得分的关系,构建基于概率的混淆矩阵
  • 评估显示不确定度分数与评分准确率强相关
  • 适合需要可靠评估结果的研究者和开发者

LLM-as-a-Judge 是广泛用于评估大型语言模型在各类任务中表现的方法。本文针对其评估不确定性量化难题提出新方法。由于大模型决策复杂且计算开销大,传统不确定性量化难以直接应用。我们提出一种新方法,通过交叉评估生成评分与可能评分之间的关系,基于词元概率构建混淆矩阵,推导出高或低不确定性的标签。在多个基准测试上验证表明,该方法所得不确定度分数与评估准确率呈强相关性,显著提升 LLM-as-a-Judge 评估的可靠性与一致性。

原文摘要 · Abstract (English)

LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty of LLM-as-a-Judge evaluations. While uncertainty quantification has been well-studied in other domains, applying it effectively to LLMs poses unique challenges due to their complex decision-making capabilities and computational demands. In this paper, we introduce a novel method for quantifying uncertainty designed to enhance the trustworthiness of LLM-as-a-Judge evaluations. The method quantifies uncertainty by analyzing the relationships between generated assessments and possible ratings. By cross-evaluating these relationships and constructing a confusion matrix based on token probabilities, the method derives labels of high or low uncertainty. We evaluate our method across multiple benchmarks, demonstrating a strong correlation between the accuracy of LLM evaluations and the derived uncertainty scores. Our findings suggest that this method can significantly improve the reliability and consistency of LLM-as-a-Judge evaluations.

大模型评估不确定性量化LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。