arXiv:2509.23510cs.AI2025-09EMNLP被引 1

用模型自评一致性预测其真实评分,便宜又准。

Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores

  • 让大模型自己评判自己在对决中的表现一致性。
  • 该一致性与人工评分的埃洛分相关性达91%。
  • 无需人类参与,适合快速评估新模型性能。

每天都有新的大语言模型发布,其实际表现往往与参数量不符。目前评估模型的最佳方法是通过人机对战测试获得埃洛分,但成本高昂。本文发现,当大模型被要求判断其他模型在对决中的优劣时,其选择结果的一致性与自身的人工埃洛分有91%的相关性。这一特性可作为埃洛分的低成本替代指标,无需人类标注或先验知识即可计算。

原文摘要 · Abstract (English)

New large language models (LLMs) are being released every day. Some perform significantly better or worse than expected given their parameter count. Therefore, there is a need for a method to independently evaluate models. The current best way to evaluate a model is to measure its Elo score by comparing it to other models in a series of contests - an expensive operation since humans are ideally required to compare LLM outputs. We observe that when an LLM is asked to judge such contests, the consistency with which it selects a model as the best in a matchup produces a metric that is 91% correlated with its own human-produced Elo score. This provides a simple proxy for Elo scores that can be computed cheaply, without any human data or prior knowledge.

模型评估埃洛分一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。