让模型互出题互解题,动态提升评测难度
MathDuels: A Self-Play Benchmark That Grows
- 模型轮流出题和解题,通过对抗性提示生成挑战
- 19个前沿模型显示出题与解题能力可分离
- 难度随新模型加入自动升级,适合评估进化能力
随着前沿语言模型在静态数学评测中接近性能天花板,现有评估已难以区分模型能力,主要因它们仅将模型视为固定题集的求解器。我们提出MathDuels,一个自对弈基准,模型同时承担出题与解题双重角色:每个模型在对抗性提示下生成数学题,并求解其他参与者创作的题目。题目通过三阶段生成流程(元提示、题目生成、难度增强)生成,并由独立验证器排除病态问题。采用Rasch模型联合估计求解者能力与题目难度;出题质量由作者所生成题目的难度决定。19个前沿模型的实验表明,出题与解题能力部分解耦,双角色评估揭示了单角色基准无法捕捉的能力差异。当新模型加入时,其生成的问题能击败先前主导的求解器,因此基准难度随参与者实力共同演化,而非固化于固定上限。我们提供公开排行榜,随新模型发布实时更新。
原文摘要 · Abstract (English)
As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, largely because they cast models solely as solvers of fixed problem sets. We introduce MathDuels, a self-play benchmark in which models occupy dual roles: each authors math problems under adversarial prompting and solves problems authored by every other participant. Problems are produced through a three-stage generation pipeline (meta-prompting, problem generation, and difficulty amplification), and validated by an independent verifier that excludes ill-posed questions. A Rasch model (Rasch, 1993) jointly estimates solver abilities and problem difficulties; author quality is derived from the difficulties of each model's authored problems. Experiments across 19 frontier models reveal that authoring and solving capabilities are partially decoupled, and that dual-role evaluation reveals capability separations invisible in single-role benchmarks. As newer models enter the arena, they produce problems that defeat previously dominant solvers, so the benchmark's difficulty co-evolves with participant strength rather than saturating at a fixed ceiling. We host a public leaderboard that updates as new models are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。