arXiv:2509.18658cs.CL2025-09EMNLP被引 27

用置信区间量化大模型评分不确定性,提升评估可靠性

Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction

  • 基于共形预测构建评分置信区间,单次运行即可生成
  • 区间覆盖率达标,中点得分比原始分数偏差更小
  • 适合对评估结果可信度有要求的应用场景

LLM-as-a-judge作为一种利用大语言模型评估自然语言生成的新范式,其评估结果的不确定性尚未得到充分研究,这可能限制其在实际应用中的部署。本文首次提出一个框架,通过共形预测为基于LLM的评分提供预测区间,以分析其不确定性。共形预测可从单次评估中生成连续预测区间,并针对离散评分任务设计了序数边界调整方法。同时,我们提出使用区间中点作为低偏差替代方案,优于原始模型得分和加权平均。大量实验表明,该方法能提供具有覆盖率保证的有效预测区间,并验证了区间中点与重新提示裁判的实用性。

原文摘要 · Abstract (English)

LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains underexplored. This lack of reliability may limit its deployment in many applications. This work presents the first framework to analyze the uncertainty by offering a prediction interval of LLM-based scoring via conformal prediction. Conformal prediction constructs continuous prediction intervals from a single evaluation run, and we design an ordinal boundary adjustment for discrete rating tasks. We also suggest a midpoint-based score within the interval as a low-bias alternative to raw model score and weighted average. We perform extensive experiments and analysis, which show that conformal prediction can provide valid prediction interval with coverage guarantees. We also explore the usefulness of interval midpoint and judge reprompting for better judgment.

大模型评估不确定性共形预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。