模型表现随生成预算变化,排名会反转,需按预算评估。
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

- 改变最大生成词数,发现模型排名随预算变动
- 超15%任务在更多预算下准确率反而下降
- 适合需要精准控制推理资源的研究者
大语言模型的标准评估假设其排名在不同推理条件下保持稳定。本文通过设置7种生成词数预算(64至4096),在三个推理基准上对四个模型进行评估(共56,476次推理)。主要发现:(i) 3%–19%的任务出现非单调现象(预算增加时准确率下降),且该现象具有模型特异性(跨模型重叠仅6%–14%);(ii) 所有基准上模型排名均随预算反转(p < 0.01,McNemar检验);(iii) 优化器分析显示模型互补性最高达+27.8个百分点,尤其在预算受限时显著;(iv) 一个预算感知路由器可捕获跨域14.1%的最优差距;预算特征在域内提升+1.6至+5.7个百分点,但具域特定性且损害跨域迁移(-1.2个百分点)。结果表明应采用预算相关的评估协议。
原文摘要 · Abstract (English)
Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks ($p {<} 0.01$, McNemar). (iii) Oracle analysis reveals model complementarity up to $+27.8$pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain ($+1.6$ to $+5.7$pp) but are domain-specific and hurt transfer ($-1.2$pp). These results argue for budget-conditioned evaluation protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。