用双向偏好熵提升大模型判断的可靠性与覆盖率
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
- 引入双向偏好熵生成无偏置的不确定性信号
- 在α=0.10时实测错误率约0.098,满足风险约束
- 相比基线可多接受2.4倍判断,适合高覆盖评估场景
大语言模型常用于成对评价中的自动判别,但易出现校准偏差与系统性偏见。本文提出 extsc{Scope}(选择性共形优化成对评价)框架,通过校准接受阈值,在交换性假设下确保未弃权判断的误差率不超过用户指定的α水平。为提供无偏的不确定性信号,我们引入双向偏好熵(BPE),通过正反位置响应查询并以平均偏好概率转换为基于熵的评分。在多个成对评价基准上,BPE在校准性和区分度上均优于标准置信度代理; extsc{Scope}在α=0.10时始终满足目标风险界,实测错误率(经验FDR)约为0.097–0.099,且保持较高覆盖率。相比原始基线, extsc{Scope}在相同风险约束下可接受最多2.4倍的判断,证明了BPE能实现可靠且高覆盖率的大模型评价。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose \textsc{Scope} (Selective Conformal Optimized Pairwise Evaluation), a framework that calibrates an acceptance threshold so that, under exchangeability, the error rate among non-abstained judgments is at most a user-specified level $α$. To supply \textsc{Scope} with a bias-neutral uncertainty signal, we introduce Bidirectional Preference Entropy (BPE), which queries the judge under both response positions and converts the order-averaged preference probability into an entropy-based score. Across various pairwise judging benchmarks, BPE outperforms standard confidence proxies in calibration and discrimination, while \textsc{Scope} consistently satisfies the target risk bound (empirical FDR $\approx 0.097$--$0.099$ at $α=0.10$) and retains substantial coverage. Compared to vanilla baselines, \textsc{Scope} accepts up to $2.4\times$ more judgments under the same risk constraint, demonstrating that BPE enables reliable and high-coverage LLM-based evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。