arXiv:2606.21008cs.CLcs.AI2026-06

用自评游戏衡量大模型智能,无需真实答案却能精准比高低。

The Metanym Game: An LLM Benchmark Without Ground Truth That Rises With the Models It Measures

论文配图:The Metanym Game: An LLM Benchmark Without Ground Truth That Rises With the Models It Measures
图 1 · 摘自论文原文
  • 模型互评生成的类比句,评分由模型自身判断。
  • 与专家题库相关性高达0.97,无数据泄露风险。
  • 适合评估自进化AI系统,可随模型能力动态扩展。

我们提出证据表明类比是大模型智能的核心。在该基准中,大模型竞争生成类比语句,并基于事实正确性、美感、智能度、独特性、长度和结构多样性等维度相互评分。所有内容均在游戏过程中生成,唯一输入是规则本身;评分完全来自模型自身的判断。真实答案被事实评分矩阵的SVD分解所替代,首次实现对大模型评审团的自我评估——即‘评委评价评委’的特征方程。对于美感等主观标准,通过评分一致性加权评委。最佳生成者往往是中等水平的评判者。尽管与人类专家设计的GPQA Diamond题库方法迥异,两者相关性仍达皮尔逊r=0.97(95%置信区间[0.92, 0.99]),未发现数据泄露。由前五名模型组成的评审团发布官方评分,其可争议席位使该基准可扩展至任意数量参与者,并随被测模型能力同步提升,或成为自进化AI的潜在引导信号。游戏融合至少八种智能维度,总分反映综合能力,各分项支持还原论分析。所有数据均可通过开源代码包重新计算:https://github.com/dnordfors/metanym-game-paper

原文摘要 · Abstract (English)

We present evidence that analogy is at the core of LLM intelligence. In our benchmark, LLMs compete in generating sets of analogous statements and rate each other's sets on their own understandings of factual correctness, beauty, intelligence, distinctness, length, and structural diversity. Nothing enters from outside: the only given is the game rules; every item is generated in play; the scores come from the players' ratings alone. Ground truth is replaced by the SVD of the factual rating matrix, which scores players as generators and judges at once -- to our knowledge the first eigen-equation that judges the judges for an LLM council-of-peers. For subjective criteria like beauty, judges are weighted by their rating consistency. The best generators turn out to be middling judges. GPQA Diamond -- difficult multiple-choice questions written by human experts -- could not be more different in method, yet the two benchmarks correlate at Pearson $r = 0.97$, 95% CI [0.92, 0.99]; no leakage could be found. A council of the five best issues the official ratings; its contestable seats let the benchmark scale to any number of players and rise with the models it measures -- a candidate steering signal for self-improving AI. Playing interweaves at least eight constructs of intelligence; the total scores the broad composite, the components allow reductionistic analysis. Every number recomputes from a released package at https://github.com/dnordfors/metanym-game-paper

大模型评估自评机制智能测量无真值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。