用大模型当裁判,让过时评测重新分出高低。
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

- 让大模型当元裁判,自动生成评判标准并淘汰候选结果。
- 在4个饱和评测中,准确率接近全量打分,仅需11.89次调用。
- 适合想高效比较模型性能的研究者和工程师。
现有语言模型评测集日趋饱和,前沿系统得分相近,常规指标难以区分。本文提出一种自改进的评估协议SEAL(Saturated Evaluation with Adaptive LLM-as-a-Meta-Judge),通过种子淘汰机制,结合任务级原则与自进化检查清单,从饱和评测中挖掘潜在排序信号。我们在代码生成、数学推理、知识密集型问答及工具使用代理任务等多类饱和基准上评估SEAL,结果表明其在排名准确性-延迟权衡上优于现有方法,达到与全量两两对比评分0.83–1.00的斯皮尔曼相关性,且在4项任务中均实现4/4的最高排名一致率,仅需11.89次模型调用,远低于全量对比所需的28.00次。
原文摘要 · Abstract (English)
Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than constructing harder alternatives, we ask whether existing tasks can be made informative again through improved evaluation over the same candidate outputs. Therefore, we present Seeded Elimination with Adaptive LLM-as-a-Meta-Judge, a self-improving evaluation protocol for extracting latent ranking signal from saturated benchmarks. SEAL seeds candidate outputs into a single elimination and evaluates each match with task-level principles plus self-improving checklist criteria. We evaluate SEAL on multiple saturated benchmarks covering code generation, mathematical reasoning, knowledge-intensive question answering, and tool-use agent task completion. Across these settings, SEAL improves the ranking-accuracy--latency trade-off over competing protocols, attaining 0.83--1.00 Spearman agreement with full pairwise judging and 4/4 top-1 agreement, while requiring only 11.89 calls per task compared with 28.00 for full pairwise evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。