arXiv:2505.21972cs.LGcs.AI2025-05中稿 · AISTATS 2026被引 4

用几何视角分析大模型自评的可靠性,发现二分评分更可信。

LLMs Judging LLMs: A Simplex Perspective

  • 将评分系统建模为概率单纯形,用几何关系判断排名是否可识别。
  • 实验证明仅靠大模型评分在多数数据集上有效,但多级评分易出错。
  • 提出贝叶斯先验建模评委质量不确定性,显著提升评估覆盖率。

针对大语言模型(LLMs)生成内容的自动评估难题,当前常用方法是使用其他大模型作为评判者,无需人工标注。这种方法隐含假设仅存在采样波动(随机不确定性),却忽略评判者质量的不确定性(认知不确定性)。若评判者完全准确,则此法合理,但其理论有效性与实际鲁棒性尚不明确。本文从新几何视角研究该问题:对于M级评分体系,所有评判者与候选模型均可表示为(M-1)维概率单纯形上的点,几何量(如三角形面积)对应关键排序概念。该视角揭示了排序可识别的直观理论条件,并给出形式化依据支持‘常识’——对双级评分(M=2)而言,大模型评判者比多级评分(M>2)更有效。基于单纯形,我们设计了编码认知不确定性的几何贝叶斯先验,并通过调整先验进行敏感性分析。在多个大模型基准测试中,实验表明仅依赖大模型评判者的排名在多数数据集上表现稳健,但也存在例外,凸显其广泛应用的同时需保持谨慎。我们的贝叶斯方法相比现有方法实现更高覆盖率,突显建模认知不确定性的关键作用。

原文摘要 · Abstract (English)

Given the challenge of automatically evaluating free-form outputs from large language models (LLMs), an increasingly common solution is to use LLMs themselves as the judging mechanism, without any gold-standard scores. Implicitly, this practice accounts for only sampling variability (aleatoric uncertainty) and ignores uncertainty about judge quality (epistemic uncertainty). While this is justified if judges are perfectly accurate, it is unclear when such an approach is theoretically valid and practically robust. We study these questions for the task of ranking LLM candidates from a novel geometric perspective: for $M$-level scoring systems, both LLM judges and candidates can be represented as points on an $(M-1)$-dimensional probability simplex, where geometric concepts (e.g., triangle areas) correspond to key ranking concepts. This perspective yields intuitive theoretical conditions and visual proofs for when rankings are identifiable; for instance, we provide a formal basis for the ``folk wisdom'' that LLM judges are more effective for two-level scoring ($M=2$) than multi-level scoring ($M>2$). Leveraging the simplex, we design geometric Bayesian priors that encode epistemic uncertainty about judge quality and vary the priors to conduct sensitivity analyses. Experiments on LLM benchmarks show that rankings based solely on LLM judges are robust in many but not all datasets, underscoring both their widespread success and the need for caution. Our Bayesian method achieves substantially higher coverage rates than existing procedures, highlighting the importance of modeling epistemic uncertainty.

大模型评估几何建模贝叶斯方法认知不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。