arXiv:2504.20879cs.AIcs.CL2025-04NeurIPS被引 67

Chatbot Arena排行榜存在数据偏见,导致部分模型因私测和高权重而虚高评分。

The Leaderboard Illusion

  • 允许厂商私测多版本并选择最佳成绩上报,造成评分偏差。
  • 闭源模型获19.2%与20.4%的评测数据,远超83个开源模型共有的29.7%。
  • 私测优势可带来最高112%的性能提升,加剧对排行榜的过拟合。

衡量进展是科学发展的基础,但随着基准测试日益重要,其也愈发易受扭曲。Chatbot Arena已成为评估顶尖AI系统的主流排行榜。本文揭示其系统性问题:未公开的私测机制使少数厂商可在发布前反复测试多个变体,并选择最优得分上报,导致评分被人为抬高。我们发现Meta在Llama-4发布前测试了27个私有LLM变体。此外,专有闭源模型被抽样频率更高(战斗次数更多),且更少被移除,形成数据获取上的巨大不平等。谷歌和OpenAI分别获得约19.2%和20.4%的总评测数据,而83个开源模型合计仅占29.7%。我们的保守估计显示,额外获得少量数据即可带来高达112%的排行榜表现提升。这些机制导致模型更倾向于过拟合特定评测环境,而非提升通用能力。尽管平台依赖组织者与社区的共同努力,我们提出具体改进建议,以推动更公平、透明的基准评估体系。

原文摘要 · Abstract (English)

Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. We show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on the arena distribution, based on our conservative estimates. Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field

AI评估排行榜偏见大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。