让弱模型通过挑战强模型来区分其能力差异。
SKATE, a Scalable Tournament Eval: Weaker LLMs differentiate between stronger ones using verifiable challenges
- 模型互设可验证问题,自动评估彼此能力。
- 6个前沿大模型中弱模型能准确识别强模型差异。
- 无需人工参与,适合快速迭代的模型评估。
评估基础模型的能力与风险至关重要,但现有方法依赖大量领域知识,难以随模型快速演进而扩展。我们提出SKATE:一种新型评估框架,让大语言模型(LLMs)相互生成并解答可验证任务。核心思想是将评估视为博弈:模型既是出题者又是答题者,激励其设计能凸显自身优势、暴露他人弱点的问题。SKATE实现全自动、无数据、可扩展,无需人类输入或专业知识。通过可验证任务替代LLM裁判,评分客观。相比局限于特定领域的程序化基准(如国际象棋或空间推理),由LLM创造性出题实现开放且可扩展的评估。作为概念验证,我们引入由LLM设定的代码输出预测(COP)挑战作为可扩展框架。采用TrueSkill排名系统评估6个前沿模型,发现:(1) 弱模型可可靠区分并评分更强模型;(2) LLM系统具备自我偏好行为,生成与其自身能力一致的问题;(3) SKATE自动揭示模型间细微能力差异。这些发现为应对大模型快速进展的通用、可扩展评估框架迈出重要一步。
原文摘要 · Abstract (English)
Evaluating the capabilities and risks of foundation models is paramount, yet current methods demand extensive domain expertise, hindering their scalability as these models rapidly evolve. We introduce SKATE: a novel evaluation framework in which large language models (LLMs) compete by generating and solving verifiable tasks for one another. Our core insight is to treat evaluation as a game: models act as both task-setters and solvers, incentivized to create questions which highlight their own strengths while exposing others' weaknesses. SKATE offers several key advantages, balancing scalability, open-endedness, and objectivity. It is fully automated, data-free, and scalable, requiring no human input or domain expertise. By using verifiable tasks rather than LLM judges, scoring is objective. Unlike domain-limited programmatically-generated benchmarks (e.g. chess-playing or spatial reasoning), having LLMs creatively pose challenges enables open-ended and scalable evaluation. As a proof of concept, we introduce LLM-set code-output-prediction (COP) challenges as a verifiable and extensible framework in which to test our approach. Using a TrueSkill-based ranking system, we evaluate six frontier LLMs and find that: (1) weaker models can reliably differentiate and score stronger ones, (2) LLM-based systems are capable of self-preferencing behavior, generating questions that align with their own capabilities, and (3) SKATE automatically surfaces fine-grained capability differences between models. Our findings are an important step towards general, scalable evaluation frameworks which can keep pace with LLM progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。