arXiv:2602.16610cs.CLcs.AI2026-02中稿 · ICML被引 6

让大模型当陪审团,自动评估生成文本质量并识别不可靠评委。

Who can we trust? LLM-as-a-jury for Comparative Assessment

  • 用改进的布拉德利-特瑞模型,从对比判断中同时推断内容排名和评委可信度。
  • 在多个评测数据集上表现优于平均法,且学习到的可信度与判别一致性高度相关。
  • 无需人工标注即可无监督校准评委,适合缺乏标注数据的自动评估场景。

大语言模型(LLMs)正被广泛用于自然语言生成的自动评估,通常采用成对比较判断。现有方法多依赖单一评委或简单聚合,假设所有评委可靠性相同。然而实践中,不同模型在不同任务和评估维度上的表现差异显著,其判断概率可能存在偏差与不一致。此外,评委校准所需的人工标注可能不可用。我们首次实证发现,大模型在比较判断中的概率存在不一致性,并表明这会限制基于概率的直接排序效果。为此,我们研究了‘大模型作为陪审团’的设定,提出 BT-sigma——一种引入每个评委判别参数的布拉德利-特瑞模型扩展,仅通过成对比较即可联合推断项目排名与评委可靠性。在基准NLG评估数据集上的实验显示,BT-sigma 持续优于基于平均的聚合方法,且学习到的判别器与独立测量的循环一致性高度相关。进一步分析表明,该方法可视为一种无监督校准机制,通过建模评委可靠性提升评估聚合效果。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly applied as automatic evaluators for natural language generation assessment often using pairwise comparative judgements. Existing approaches typically rely on single judges or aggregate multiple judges assuming equal reliability. In practice, LLM judges vary substantially in performance across tasks and evaluation aspects, and their judgment probabilities may be biased and inconsistent. Furthermore, human-labelled supervision for judge calibration may be unavailable. We first empirically demonstrate that inconsistencies in LLM comparison probabilities exist and show that it limits the effectiveness of direct probability-based ranking. To address this, we study the LLM-asa-jury setting and propose BT-sigma, a judge-aware extension of the Bradley-Terry model that introduces a discriminator parameter for each judge to jointly infer item rankings and judge reliability from pairwise comparisons alone. Experiments on benchmark NLG evaluation datasets show that BT-sigma consistently outperforms averaging-based aggregation methods, and that the learned discriminators strongly correlate with independent measures of the cycle consistency of LLM judgments. Further analysis reveals that BT-sigma can be interpreted as an unsupervised calibration mechanism that improves aggregation by modelling judge reliability.

大模型评估陪审团模型无监督校准生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。