用伊辛模型捕捉大模型评委间的依赖关系,提升评分准确性
Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models
- 引入伊辛图模型建模大模型评委间依赖关系
- 在真实数据集上相比传统方法准确率显著提升
- 适合评估大模型输出质量的研究者使用
大规模AI评估越来越多地依赖于从K名标注者(包括作为评委的大模型)获取二值判断进行聚合。经典方法如Dawid-Skene或加权多数投票假设标注者在真实标签Y∈{0,1}给定下条件独立,这一假设常被大模型评委违反,因其共享数据、架构、提示词和失败模式。忽略此类依赖会导致后验校准错误,甚至产生自信的错误预测。我们研究基于伊辛图模型和潜在因子的依赖感知标签聚合层次模型。对于类别相关伊辛模型,贝叶斯对数几率通常在投票数上为二次型;对于类别无关耦合,简化为带相关性调整参数的线性加权投票。我们给出有限K的例子,显示基于条件独立的方法即使匹配单个标注者边缘分布,仍可能反转贝叶斯标签。我们证明分离结果表明,随着裁判数量增加,这些方法始终严格次优,在潜在因子下产生非消失的超额风险。最后,我们在三个真实世界数据集上评估所提方法,展示其优于经典基线。
原文摘要 · Abstract (English)
Large-scale AI evaluation increasingly relies on aggregating binary judgments from $K$ annotators, including LLMs used as judges. Most classical methods, e.g., Dawid-Skene or (weighted) majority voting, assume annotators are conditionally independent given the true label $Y\in\{0,1\}$, an assumption often violated by LLM judges due to shared data, architectures, prompts, and failure modes. Ignoring such dependencies can yield miscalibrated posteriors and even confidently incorrect predictions. We study label aggregation through a hierarchy of dependence-aware models based on Ising graphical models and latent factors. For class-dependent Ising models, the Bayes log-odds is generally quadratic in votes; for class-independent couplings, it reduces to a linear weighted vote with correlation-adjusted parameters. We present finite-$K$ examples showing that methods based on conditional independence can flip the Bayes label despite matching per-annotator marginals. We prove separation results demonstrating that these methods remain strictly suboptimal as the number of judges grows, incurring nonvanishing excess risk under latent factors. Finally, we evaluate the proposed method on three real-world datasets, demonstrating improved performance over the classical baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。