LLM当裁判常自大,新方法让判断更靠谱
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
- 用集成框架把LLM变成能评估自身可信度的裁判
- 发现当前LLM判断自信程度远超实际正确率
- 适合需要可靠、可调风险评估的AI评测场景
大型语言模型(LLMs)被广泛用于自动化评判,其实际价值取决于准确性和可信的风险感知能力。现有方法多关注准确性,忽视了良好校准的置信度对自适应、可靠评价流程的重要性。本文倡导从以准确性为中心的评估转向以置信度驱动、风险感知的LLM-as-a-Judge系统,强调校准置信度对可信和自适应评价的必要性。我们系统识别出当前LLM裁判中的过自信现象:预测置信度显著高于实际正确率,损害实际部署的可靠性。为量化该现象,我们提出TH-Score,一种衡量置信度与准确度对齐程度的新指标。此外,我们提出LLM-as-a-Fuser,一个将LLM转化为可靠、风险感知评估者的集成框架。大量实验表明,该方法显著提升校准度,实现自适应的置信度驱动评价流程,在准确性和可靠性上均优于现有基线。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are widely used as automated judges, where practical value depends on both accuracy and trustworthy, risk-aware judgments. Existing approaches predominantly focus on accuracy, overlooking the necessity of well-calibrated confidence, which is vital for adaptive and reliable evaluation pipelines. In this work, we advocate a shift from accuracy-centric evaluation to confidence-driven, risk-aware LLM-as-a-Judge systems, emphasizing the necessity of well-calibrated confidence for trustworthy and adaptive evaluation. We systematically identify the Overconfidence Phenomenon in current LLM-as-a-Judges, where predicted confidence significantly overstates actual correctness, undermining reliability in practical deployment. To quantify this phenomenon, we introduce TH-Score, a novel metric measuring confidence-accuracy alignment. Furthermore, we propose LLM-as-a-Fuser, an ensemble framework that transforms LLMs into reliable, risk-aware evaluators. Extensive experiments demonstrate that our approach substantially improves calibration and enables adaptive, confidence-driven evaluation pipelines, achieving superior reliability and accuracy compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。