用线性探针快速获得可靠不确定性估计,无需额外训练。
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- 用贝叶斯分数损失训练线性探针,从推理模型隐藏状态中提取不确定性。
- 校准效果优于现有方法,计算量减少约10倍,高置信度预测更准确。
- 适合对误报率敏感的生产环境,可直接接入现有LLM评判系统。
随着基于大语言模型的评判系统在工业应用中日益重要,高效获取可靠的不确定性估计成为部署关键。现有方法如口头置信度和多生成策略往往校准不佳或计算开销大。本文提出一种基于贝叶斯损失训练的线性探针,从推理模型的隐藏状态中直接生成校准后的不确定性估计,无需额外训练。我们在客观任务(推理、数学、事实性、编码)和主观人类偏好判断上评估该方法。结果表明,该探针在校准性上显著优于现有方法,计算成本降低约10倍,对未见评估领域具有强泛化能力,且在高置信度预测上表现更优。然而,探针给出的估计偏保守,在简单数据集上表现较差,但更适合对低假阳性率要求高的安全关键场景。整体而言,基于可解释性的不确定性估计为生产环境中大语言模型评判提供了一种实用、可扩展的即插即用方案。
原文摘要 · Abstract (English)
As LLM-based judges become integral to industry applications, obtaining well-calibrated uncertainty estimates efficiently has become critical for production deployment. However, existing techniques, such as verbalized confidence and multi-generation methods, are often either poorly calibrated or computationally expensive. We introduce linear probes trained with a Brier score-based loss to provide calibrated uncertainty estimates from reasoning judges' hidden states, requiring no additional model training. We evaluate our approach on both objective tasks (reasoning, mathematics, factuality, coding) and subjective human preference judgments. Our results demonstrate that probes achieve superior calibration compared to existing methods with $\approx10$x computational savings, generalize robustly to unseen evaluation domains, and deliver higher accuracy on high-confidence predictions. However, probes produce conservative estimates that underperform on easier datasets but may benefit safety-critical deployments prioritizing low false-positive rates. Overall, our work demonstrates that interpretability-based uncertainty estimation provides a practical and scalable plug-and-play solution for LLM judges in production.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。