用模型生成精准难度的问题,来测试和训练模型自身能力。
Ask-E: An Environment for Calibrated Question Generation

- 以两个现有模型的能力为边界,生成恰好只有其中一个能解的问题。
- 顶尖模型在基准上校准准确率不足50%,说明仍有巨大提升空间。
- 无需额外数据或奖励,训练后可提升多个数学任务表现。
当前,我们通过在模型能力前沿的问题上训练与评估来提升模型。但构造这类问题极具挑战性,需精准把控难度并超越现有问题分布。这要求理解问题求解的本质,而生成与模型能力精确匹配的问题,本身需要超越该模型的能力——随着模型进步,这一约束愈发沉重。我们的核心洞察是:能够持续生成校准问题的模型,必然具备超越目标能力。因此,我们提出 Ask-E 环境,用于评估和训练模型生成特定技能水平问题的能力,而非回答问题。具体而言,目标技能水平定义为两个现有语言模型能力范围之间的区间。若一个问题恰好仅被其中一模型解决,则视为成功校准,精准定位在目标范围内,并区分两模型能力差异。Ask-E 既作为基准,也作为训练环境,支持模型生成多种难度的问题。实验发现,即使前沿模型在该基准上校准准确率仍低于50%,表明未来仍有显著进步空间。此外,仅在该环境训练即可在多个下游数学基准上取得提升,且无需新数学数据、不与更强模型交互、无正确性奖励。
原文摘要 · Abstract (English)
Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to probe model limits and generalize beyond existing question distributions. It also means placing problems at a precise difficulty level, which requires understanding what it takes to solve them. In short, generating problems calibrated to a model's current frontier demands capability beyond it, an increasingly burdensome constraint as models improve. Our key insight is that we can leverage this constraint to our advantage: a model that can generate problems consistently calibrated to a given frontier must possess capability beyond it. Accordingly, we present Ask-E, an environment that benchmarks and trains models on their ability to write questions at a given skill level, rather than answer them. Concretely, we define target skill levels as ranges bounded by the capabilities of two existing language models. A generated question is successfully calibrated if exactly one of the two models can solve it, placing it precisely within the target range and differentiating the capabilities of these models. Ask-E serves both as a benchmark and a training environment, where models generate problems calibrated to a variety of skill levels. We find that even frontier models achieve below 50% calibration on the benchmark, leaving significant headroom to measure future progress. We also show that training on this environment leads to improvements across a number of downstream math benchmarks even with no new math data, no interaction with stronger models, and no correctness-based reward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。