arXiv:2606.17930cs.AI2026-06被引 2

测试不同推理计算资源下大模型表现,发现评测结果高度依赖计算设置。

How Inference Compute Shapes Frontier LLM Evaluation

论文配图:How Inference Compute Shapes Frontier LLM Evaluation
图 1 · 摘自论文原文
  • 通过扩大上下文长度、重复提交等方法控制推理算力,评估模型表现。
  • 更大计算预算显著提升模型在多个领域的得分,最高提升达数倍。
  • 评测应公开计算协议,避免单一预算误导模型真实能力判断。

AI评估正转向更复杂的任务,这些任务需要更长的推理轨迹和工具使用。因此,模型表现越来越依赖测试时可用的计算资源(即推理算力)。然而,许多评估仅在单一受限预算下报告性能,导致低分可能反映的是评估设定而非模型本身的能力。为此,我们在涵盖软件工程、数学、医学和网络安全的七个挑战性基准上,评估了最多12个前沿语言模型。采用三种简单的推理扩展干预措施:更大的令牌预算、上下文压缩和由模型自身或最小正确性反馈引导的重复提交尝试。结果显示:第一,更大的令牌预算在多个领域(包括网络安全、FrontierMath、Humanity's Last Exam 和 TerminalBench)显著提升性能;第二,固定预算评估会逐渐低估前沿模型的真实能力,新模型在高预算下能解锁更难任务并更可靠解决;第三,不同基准对不同推理扩展方法的响应差异明显:重复提交普遍有效,但大令牌预算、外部反馈和并行尝试的价值因任务而异。总体而言,基准得分具有协议依赖性。我们主张评测应报告模型能力随推理算力的变化,明确说明协议选择,并在共享的大算力范围内进行跨代模型比较,尤其在安全或政策相关场景中。

原文摘要 · Abstract (English)

AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.

大模型评估推理算力评测协议多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。