arXiv:2507.22359cs.AIcs.CL2025-07ACL

让大模型互相打分,不用基准测试也能看清谁更强。

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models

  • 大模型组队互评,动态透明无偏见。
  • 排名稳定度达70.7%,能区分模型真实能力。
  • 发现记忆答题现象,适合评估研究者用。

尽管大语言模型在众多任务中表现出色,但可靠评估仍面临数据污染、机制不透明和主观偏好等挑战。为此,我们提出「大模型联赛」(LOL),一种无需基准的互评范式,将多个大模型组织为自管理联盟,进行多轮相互评价。LOL融合动态、透明、客观、专业四重标准,有效克服现有方法缺陷。在数学与编程任务上对八款主流大模型的实验表明,该方法可准确区分模型能力,内部排名稳定性高(顶k一致性达70.7%)。除排名外,还揭示了传统方法难以捕捉的实证现象,如部分模型存在“基于记忆的回答”行为,且OpenAI家族模型内部得分更高(Δ=9,p<0.05)。我们已开源框架与代码,作为当前评估体系的重要补充。

原文摘要 · Abstract (English)

Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination, opaque operation, and subjective preferences. To address these issues, we propose League of LLMs (LOL), a novel benchmark-free evaluation paradigm that organizes multiple LLMs into a self-governed league for multi-round mutual evaluation. LOL integrates four core criteria (dynamic, transparent, objective, and professional) to mitigate key limitations of existing paradigms. Experiments on eight mainstream LLMs in mathematics and programming demonstrate that LOL can effectively distinguish LLM capabilities while maintaining high internal ranking stability (Top-$k$ consistency $= 70.7\%$). Beyond ranking, LOL reveals empirical findings that are difficult for traditional paradigms to capture. For instance, ``memorization-based answering'' behaviors are observed in some models, and higher in-family scores are found in the OpenAI model family ($Δ= 9$, $p < 0.05$). Finally, we make our framework and code publicly available as a valuable complement to the current LLM evaluation ecosystem.

大模型评估互评机制无基准排名稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。