研究基准测试如何诱使模型开发者投机取巧,提出新评估方式可让排名真正反映模型真实水平。
Leaderboard Incentives: Model Rankings under Strategic Post-Training
- 将基准测试建模为设计者与开发者间的博弈,分析激励机制
- 现有基准导致无稳定策略平衡,引发不可预测的投机行为
- 新协议「先调优再测试」可实现唯一稳定均衡,按真实质量排序
主流基准测试促使模型开发者在训练后刻意优化榜单表现(称为benchmaxxing),形成策略性资源分配。本文首次从博弈论角度系统分析这一激励结构:将基准设计视为领导者-跟随者博弈,其中设计者制定评估协议,多个开发者在给定协议下同时竞争。每个开发者拥有未知潜质的模型,可通过投入资源提升其在特定基准上的观测得分。我们证明,当前基准对应的博弈中不存在纳什均衡,解释了为何实际中出现激励错配与不透明策略。但若满足温和条件,近期提出的「tune-before-test」协议可构造出具有唯一纳什均衡的基准,使得排名准确反映模型潜质。结果表明,良好设计的基准能避免不良激励。
原文摘要 · Abstract (English)
Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmaxxing or training on the test task. In this work, we initiate a principled study of the incentive structure that benchmarks induce. We model benchmarking as a Stackelberg game between a benchmark designer who chooses an evaluation protocol and multiple model developers who compete simultaneously in a subgame given by the designer's choice. Each competitor has a model of unknown latent quality and can inflate its observed score by allocating resources to benchmark-specific improvements. First, we prove that current benchmarks induce games for which no Nash equilibrium between model developers exists. This result suggests one explanation for why current practice leads to misaligned incentives, prompting model developers to strategize in opaque ways. However, we prove that under mild conditions, a recently proposed evaluation protocol, called tune-before-test, induces a benchmark with a unique Nash equilibrium that ranks models by latent quality. This positive result demonstrates that benchmarks need not set bad incentives, even if current evaluations do.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。