改进大模型评判的不确定性估计,用更少比较实现更高效率。
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
- 提出广义概率建模框架,统一多种专家模型
- 不确定性估计使比较次数减少约50%,性能仍强
- 可识别低质量预测,适合需要可靠排序的场景
本文研究对比式大模型评判框架中的广义概率建模与不确定性估计。我们表明现有Product-of-Experts方法是更广泛框架的特例,支持多样建模选择。同时提出改进的单次比较不确定性估计,提升评估效率,并引入整体排名不确定性估计方法。实验显示,具体专家模型对最终排名影响有限,但所提不确定性估计(尤其是重排概率)显著提升系统效率,减少约50%的比较次数。此外,排名级不确定性指标可用于识别低质量预测,且概率模型设计对整体不确定性质量有显著影响。
原文摘要 · Abstract (English)
This paper explores generalised probabilistic modelling and uncertainty estimation in comparative LLM-as-a-judge frameworks. We show that existing Product-of-Experts methods are specific cases of a broader framework, enabling diverse modelling options. Furthermore, we propose improved uncertainty estimates for individual comparisons, enabling more efficient selection and achieving strong performance with fewer evaluations. We also introduce a method for estimating overall ranking uncertainty. Finally, we demonstrate that combining absolute and comparative scoring improves performance. Experiments show that the specific expert model has a limited impact on final rankings but our proposed uncertainty estimates, especially the probability of reordering, significantly improve the efficiency of systems reducing the number of needed comparisons by ~50%. Furthermore, ranking-level uncertainty metrics can be used to identify low-performing predictions, where the nature of the probabilistic model has a notable impact on the quality of the overall uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。