arXiv:2511.21140cs.LGcs.CL2025-11被引 26

用校准数据修正大模型评分偏差,提升评估可靠性

How to Correctly Report LLM-as-a-Judge Evaluations

  • 引入校准集构建考虑双重不确定性的置信区间
  • 在真实评分与模型判别能力特定条件下优于人工评估
  • 支持分布偏移下无偏评估,适合严谨的模型评测场景

大语言模型(LLMs)被广泛用于替代人工标注者,对模型输出进行可扩展的评估。然而,LLM评判者的敏感性与特异性不完美,会导致原始评估分数产生偏差。本文提出一种简单可插拔的校准框架,能够纠正此类偏差,并实现统计上合理的不确定性量化。该框架构建的置信区间同时考虑测试集和人工标注校准集带来的不确定性,并采用自适应策略分配校准样本以获得更紧的区间。更重要的是,我们刻画了由真实评估分数及模型判别能力决定的参数区域,在这些区域内,基于LLM的评估比纯人工评估更可靠。此外,与现有方法不同,本框架在测试集与校准集存在分布偏移时仍保持无偏性。

原文摘要 · Abstract (English)

Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensitivity and specificity of the LLM judges induce bias in naive evaluation scores. We propose a simple plug-in framework that corrects this bias and enables statistically principled uncertainty quantification. Our framework constructs confidence intervals that account for uncertainty from both the test dataset and a human-labeled calibration dataset. Additionally, it uses an adaptive strategy to allocate calibration samples for tighter intervals. Importantly, we characterize parameter regimes defined by the true evaluation score and the LLM judge's sensitivity and specificity in which our LLM-based evaluation yields more reliable estimates than human-only evaluation. Moreover, we show that our framework remains unbiased under distribution shift between the test and calibration datasets, in contrast to existing approaches.

大模型评估偏差校正置信区间分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。