提出有限校准地图,判断大模型评委组该用简单聚合还是复杂表结构。
A Finite-Calibration Regime Map for LLM Judge Panels

- 构建有限校准范式地图,自动选择评委路径与聚合方式。
- 在20个数据-预算组合中,16个场景下简单聚合误差更低。
- 适合评估大模型评分系统部署的工程人员参考。
部署大模型评委组需消耗人力标注来拟合校准器、构建候选评委路径,并验证部署方案。本文研究在有限标注下,何时应采用低维堆叠器或可靠性模型,何时值得承担无限制联合输出表的单元数量和未见模式成本。我们将此问题转化为有限校准范式地图,并实例化为有限校准评委组选择(FCPS),一个支持诊断的验证选择器,可决策评委路径、部署面板规模及聚合族。在RewardBench、LLMBar、SummEval和Arena100K上,基于七评委池,标量/可靠性聚合在20个真实数据-预算单元中有16个点估计的均方误差低于无约束联合表;11个单元的95%置信区间排除零值。更丰富的回退/收缩表缩小部分差距,但无法突破有限支持瓶颈。受控校准增长数据显示相反情形:当标注包含六重交互时,所选表格扩展至含交互前缀,一旦未见分布消失,其均方误差从0.224降至0.061。实际部署关键在于:现有标注是否足以估算下一评委的信息。
原文摘要 · Abstract (English)
Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels should support a low-dimensional stacker or reliability model, and when an unrestricted joint output table is worth its cell-count and unseen-pattern cost. We cast this as a finite-calibration regime map and instantiate it as Finite-Calibration Panel Selection (FCPS), a validation selector over judge path, deployed panel size, and aggregator family with support diagnostics. Across RewardBench, LLMBar, SummEval, and Arena100K with a seven-judge pool, scalar/reliability aggregation has lower MSE than unrestricted joint-table calibration in 16 of 20 real dataset--budget cells by point estimate, while paired 95% intervals exclude zero in 11 cells; richer backoff/shrinkage tables narrow some gaps while preserving the finite-support bottleneck. Controlled calibration-growth data show the opposite regime: when labels contain a six-way interaction, the selected table grows to the interaction-bearing prefix and its MSE falls from 0.224 to 0.061 once unseen mass vanishes. The practical deployment question is whether the next judge's information is estimable under the available human labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。