arXiv:2603.10960cs.LGmath.ST2026-03ACL被引 2

提出新方法精准比较大模型推理能力,支持高/低预算测试场景。

Ranking Reasoning LLMs under Test-Time Scaling

  • 用统计方法融合多次采样输出,实现更可靠的模型排序
  • 80次采样下与最优基准一致率超93%,19至34种方法完全匹配排序
  • 开源工具Scorio适用于科研与工程落地的模型评估

测试时扩展通过每提示多次采样评估推理大模型,但在此框架下的模型排序研究仍不充分。本文形式化了密集基准下的测试时扩展排序问题,并推出Scorio库,集成配对比较模型、项目反应理论(IRT)、投票规则、图与谱方法等统计排序技术。在四个奥数风格数学基准(AIME'24, AIME'25, HMMT'25, BrUMO'25)上评估20个推理模型,最多达80次试验,多数全试验排名与贝叶斯黄金标准$→{\mathrm{Bayes}_{\mathcal{U}}}@80$高度一致(平均肯德尔τ_b = 0.93–0.95),19–34种方法完全恢复相同排序。单次采样情形下最佳方法达到τ_b ≈ 0.86。使用贪婪解码作为经验先验($→{\mathrm{Bayes}_{\mathbf{R}_0}}@N$)可使N=1时方差降低16%–52%,但在贪婪与随机采样不一致时可能引入偏差。结果确定了高低预算测试时扩展下的可靠排序方法。Scorio已开源,地址:https://github.com/mohsenhariri/scorio。

原文摘要 · Abstract (English)

Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking models in this regime remains underexplored. We formalize dense benchmark ranking under test-time scaling and introduce Scorio, a library that implements statistical ranking methods such as paired-comparison models, item response theory (IRT) models, voting rules, and graph- and spectral-based methods. Across $20$ reasoning models on four Olympiad-style math benchmarks (AIME'24, AIME'25, HMMT'25, and BrUMO'25; up to $N=80$ trials), most full-trial rankings agree closely with the Bayesian gold standard $\mathrm{Bayes}_{\mathcal{U}}@80$ (mean Kendall's $τ_b = 0.93$--$0.95$), and $19$--$34$ methods recover exactly the same ordering. In the single-trial regime, the best methods reach $τ_b \approx 0.86$. Using greedy decoding as an empirical prior ($\mathrm{Bayes}_{\mathbf{R}_0}@N$) reduces variance at $N=1$ by $16$--$52\%$, but can bias rankings when greedy and stochastic sampling disagree. These results identify reliable ranking methods for both high- and low-budget test-time scaling. We release Scorio as an open-source library at https://github.com/mohsenhariri/scorio.

大模型评估推理能力统计排序开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。