多专家协同提升大模型对话评价的可靠性与覆盖率。
Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue

- 用评分平均与决策投票融合多专家判断,提升评估稳定性。
- 提出MC3方法,适配不同模型的评分尺度,保持风险可控。
- 在三个对话数据集上验证,显著提升异构专家场景下的表现。
本文研究用于开放式对话中成对大模型评判的多专家共形风险控制(CRC)算法。核心洞察是:多专家聚合可从源头净化评分函数,弥补单一阈值控制的不足。为此,我们设计了两种方法:评分平均与决策投票,分别在评分和决策层面进行聚合。在同质专家组中,二者均优于单专家方法;但在异构大模型评判者中,因统一阈值难以匹配各异的评分尺度,覆盖范围仍受限。为解决此问题,我们提出边际校准共形共识(MC3):通过初始阈值比例捕捉各专家的独立评分尺度,并联合优化一个统一的决策函数 $C_t(x)$,在校准与测试阶段一致应用,从而维持交换性。为评估框架,我们构建了Panel基准,包含1,800对人类成对偏好数据,基于四个开源大模型在三个领域(ESConv、MSC、DREAM)的对话上下文生成,提供完整logit访问。实验表明,评分平均与决策投票在同质面板中显著提升准确率与接受率;而MC3则在所有三个数据集上实现对异构面板的扩展性能,有效适配各专家的评分尺度。
原文摘要 · Abstract (English)
In this paper, we explore multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Our core insight is that multi-expert aggregation offers a complementary remedy to CRC: whereas CRC controls risk at the decision threshold through abstention, aggregation sanitizes the scoring function at its source. Guided by this, we first design two multi-expert CRC methods: Score Averaging and Decision Voting, which aggregate at the score and decision levels, respectively. While both strategies outperform single-expert methods on homogeneous expert panels, on heterogeneous LLM judges they remain risk-valid but recover only limited coverage, because a uniform threshold cannot match the experts' distinct scoring scales. To resolve this issue, we further propose Marginal-Calibrated Conformal Consensus (MC3): it captures distinct per-expert scales via initial threshold ratios, while jointly tuning a unified decision function $C_t(x)$ applied identically in both calibration and test, thereby preserving exchangeability. To evaluate our framework, we construct Panel, a 1,800-pair human pairwise-preference benchmark for open-ended dialogue. It is built on responses generated by four open-weight LLMs over dialogue contexts from three domains (ESConv, MSC, DREAM), with full logit access. In experiments, we find that both Score Averaging and Decision Voting substantially improve accuracy and acceptance rate on homogeneous panels. Notably, MC3 extends these gains to heterogeneous panels by accommodating distinct per-expert scoring scales across all three datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。