arXiv:2510.02306cs.CL2025-10被引 1

重新审视大模型评估中平局的含义,发现平局反映问题难度而非模型等价。

Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation

  • 将平局视为问题难度信号,而非模型实力相等
  • 忽略平局更新可提升预测准确率1-3%
  • 适合关注评估公平性与模型排名可信度的研究者

在大语言模型的竞技场式评估中,两个模型对用户问题生成回复,用户选择胜者或判定为平局,据此调整两模型评分。当前普遍做法是将对决类比为国际象棋对弈,采用埃洛评分系统及其衍生方法。本文对此范式提出质疑:平局是否真意味着两模型实力相当?我们推测,平局更可能反映问题难度——若问题过于简单,两模型均易正确回答。在三个真实世界竞技场数据集上,我们发现忽略平局评分更新可使四类评分系统的对决结果预测准确率相对提升1%-3%。进一步分析表明,平局更常出现在被评作极简单或高度客观的问题上,风险比分别为1.37和1.35。建议未来评分系统重新考虑平局语义,并在评分更新中纳入问题特性。

原文摘要 · Abstract (English)

In arena-style evaluation of large language models (LLMs), two LLMs respond to a user query, and the user chooses the winning response or deems the "battle" a draw, resulting in an adjustment to the ratings of both models. The prevailing approach for modeling these rating dynamics is to view battles as two-player game matches, as in chess, and apply the Elo rating system and its derivatives. In this paper, we critically examine this paradigm. Specifically, we question whether a draw genuinely means that the two models are equal and hence whether their ratings should be equalized. Instead, we conjecture that draws are more indicative of query difficulty: if the query is too easy, then both models are more likely to succeed equally. On three real-world arena datasets, we show that ignoring rating updates for draws yields a 1-3% relative increase in battle outcome prediction accuracy (which includes draws) for all four rating systems studied. Further analyses suggest that draws occur more for queries rated as very easy and those as highly objective, with risk ratios of 1.37 and 1.35, respectively. We recommend future rating systems to reconsider existing draw semantics and to account for query properties in rating updates.

大模型评估评分系统平局语义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。