Elo评级在非理想条件下仍表现优异,因其具备在线学习的鲁棒性。
Is Elo Rating Reliable? A Study Under Model Misspecification
- 将Elo重解释为在线梯度下降,具无遗憾保证
- 真实数据稀疏性使复杂模型反而不如Elo预测准确
- 评级精度与排名效果高度相关,适合实战场景
Elo评级广泛应用于从竞技游戏到大语言模型的技能评估,通常被视为对平稳伯莱德-特里(BT)模型的增量更新算法。然而,我们对实际匹配数据集的实证分析发现两个意外现象:(1) 多数比赛显著偏离BT模型和平稳性假设,质疑Elo的可靠性;(2) 尽管存在偏差,Elo仍经常优于更复杂的评级系统(如mElo和成对模型),尤其是在胜率预测方面,这些系统专门设计用于处理数据中的非BT成分。本文从三个关键视角解释这一反直觉现象:(a) 将Elo重新诠释为在线梯度下降,即使在模型误设和非平稳环境中也具有无遗憾保证;(b) 在来自传递但非BT模型(如强或弱随机传递模型)生成的大量合成数据上进行实验,表明实际匹配数据的稀疏性是Elo相比复杂模型在预测上表现更优的关键因素;(c) 观察到Elo预测准确率与其排名性能高度相关,进一步支持其在排序任务中的有效性。
原文摘要 · Abstract (English)
Elo rating, widely used for skill assessment across diverse domains ranging from competitive games to large language models, is often understood as an incremental update algorithm for estimating a stationary Bradley-Terry (BT) model. However, our empirical analysis of practical matching datasets reveals two surprising findings: (1) Most games deviate significantly from the assumptions of the BT model and stationarity, raising questions on the reliability of Elo. (2) Despite these deviations, Elo frequently outperforms more complex rating systems, such as mElo and pairwise models, which are specifically designed to account for non-BT components in the data, particularly in terms of win rate prediction. This paper explains this unexpected phenomenon through three key perspectives: (a) We reinterpret Elo as an instance of online gradient descent, which provides no-regret guarantees even in misspecified and non-stationary settings. (b) Through extensive synthetic experiments on data generated from transitive but non-BT models, such as strongly or weakly stochastic transitive models, we show that the ''sparsity'' of practical matching data is a critical factor behind Elo's superior performance in prediction compared to more complex rating systems. (c) We observe a strong correlation between Elo's predictive accuracy and its ranking performance, further supporting its effectiveness in ranking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。