改进ELO评分体系,让大模型评估更稳定准确
am-ELO: A Stable Framework for Arena-based LLM Evaluation
- 用最大似然估计替代迭代更新,提升排名稳定性
- 引入标注者能力评估,实现模型与标注者双估
- 实验验证框架更稳健,适合大模型评估场景
基于竞技场的评估是现代人工智能模型(尤其是大语言模型)的核心评价范式。现有基于ELO评分体系的框架因排名不一致及未考虑标注者能力差异,存在固有的不稳定性问题。本文提出新型稳定竞技场评估框架am-ELO,通过将ELO的迭代更新替换为最大似然估计(MLE)方法,构建m-ELO,并提供理论证明其排名一致性与稳定性。进一步,am-ELO改进ELO的概率函数,融合标注者能力建模,实现模型分数与标注者可靠性的同时估计。实验表明,该方法显著提升评估稳定性,为大语言模型提供更可靠、准确和稳定的评估方案。
原文摘要 · Abstract (English)
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs). Existing framework based on ELO rating system suffers from the inevitable instability problem due to ranking inconsistency and the lack of attention to the varying abilities of annotators. In this paper, we introduce a novel stable arena framework to address these issues by enhancing the ELO Rating System. Specifically, we replace the iterative update method with a Maximum Likelihood Estimation (MLE) approach, m-ELO, and provide theoretical proof of the consistency and stability of the MLE approach for model ranking. Additionally, we proposed the am-ELO, which modify the Elo Rating's probability function to incorporate annotator abilities, enabling the simultaneous estimation of model scores and annotator reliability. Experiments demonstrate that this method ensures stability, proving that this framework offers a more robust, accurate, and stable evaluation method for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。