arXiv:2510.22219cs.CLcs.AI2025-10

不依赖真实答案,评估大模型在文本对比中的错误率。

Estimating the Error of Large Language Models at Pairwise Text Comparison

  • 通过两轮对比构建排名,估计模型错误率。
  • 六款模型错误率相近,位置偏差小,表现稳定。
  • 适合评估大模型决策可靠性,尤其关注提示词鲁棒性。

我们测量了大语言模型在成对文本比较中的输出错误率,分析其偏好判断的错误概率。该方法无需真实答案,支持两种场景:(i) 错误率与比较顺序无关,每对文本仅需两次对比(任一文本先出现);(ii) 假设存在二元位置偏差,即不同顺序下错误率不同,通过重复对比估算。采用柯佩兰德计数法从成对偏好中构建文本排序,揭示了基于大模型的成对比较在扩展性上的不足,并辅助估计错误率。我们在六款大模型(ChatGPT、Claude、DeepSeek、Gemini、Grok、Qwen)上测试五类文本输入,获得一致的错误率估计。总体来看,两种位置偏差项接近且趋同于均匀误差。综合错误率与提示词变化下的鲁棒性,Claude 表现最优。所提方法优于有偏的 Bradley-Terry 模型和可交换性评分,在识别模型错误方面更具优势。

原文摘要 · Abstract (English)

We measure LLMs' output error at pairwise text comparison, noting the probability of error in their preferences. Our method does not rely on the ground truth and supports two scenarios: (i) uniform error rate regardless of the order of comparison, estimated with two comparisons for each text pair with either text placed first; (ii) binary positional bias assuming distinct error rates for the two orders of comparison, estimated with repeated comparisons between the texts. The Copeland counting constructs a ranking over the compared texts from pairwise preferences; the ranking reveals the poor scalability of LLM-based pairwise comparison and helps yield the estimates for LLMs' error rates. We apply the method to six LLMs (ChatGPT, Claude, DeepSeek, Gemini, Grok, Qwen) with five types of text input and obtain consistent estimates of LLMs' error. In general, the measured two positional bias terms are similar, close to the uniform error. Considering both the error rates and the robustness to the variation of prompts, Claude obtained the most desirable performance in this experiment. Our model outperforms the biased Bradley-Terry model and the commutativity score in indicating LLMs' error at this task.

大模型评估错误率估计文本对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。