arXiv:2510.22593cs.CLcs.AI2025-10被引 1

用模型互评自动评估大模型,避免数据污染。

AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment

  • 模型轮流当出题人、答题人和裁判,动态生成任务。
  • 与主流基准相关性达78%(MMLU-Pro)和63%(GPQA)。
  • 多裁判设计更可靠,适合持续评估演化中的模型。

我们提出 AutoBench,一个完全自动化且自维持的大型语言模型(LLM)评估框架,通过相互评价实现评估。该方法最初由 eZecute S.R.L. 作为开源项目开发。与易受测试集污染且适应性差的静态基准不同,AutoBench 动态生成新任务,模型在多个领域中交替担任问题生成者、参赛者和裁判角色。通过迭代加权机制强化一贯可靠的评估者影响,将同行判断聚合为反映集体共识的排名。实验表明,该方法与 MMLU-Pro(相关性78%)和 GPQA(相关性63%)等权威基准具有强相关性,验证了其互评范式的有效性。多裁判设计显著优于单裁判基线,证明分布式评估更具鲁棒性和与人类一致性。AutoBench 提供了一种可扩展、抗污染的替代方案,适用于持续评估不断演化的语言模型。

原文摘要 · Abstract (English)

We present AutoBench, a fully automated and self-sustaining framework for evaluating Large Language Models (LLMs) through reciprocal peer assessment. This paper provides a rigorous scientific validation of the AutoBench methodology, originally developed as an open-source project by eZecute S.R.L.. Unlike static benchmarks that suffer from test-set contamination and limited adaptability, AutoBench dynamically generates novel evaluation tasks while models alternately serve as question generators, contestants, and judges across diverse domains. An iterative weighting mechanism amplifies the influence of consistently reliable evaluators, aggregating peer judgments into consensus-based rankings that reflect collective model agreement. Our experiments demonstrate strong correlations with established benchmarks including MMLU-Pro and GPQA (respectively 78\% and 63\%), validating this peer-driven evaluation paradigm. The multi-judge design significantly outperforms single-judge baselines, confirming that distributed evaluation produces more robust and human-consistent assessments. AutoBench offers a scalable, contamination-resistant alternative to static benchmarks for the continuous evaluation of evolving language models.

模型评估互评机制LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。