arXiv:2512.02008cs.CL2025-12被引 11

如何在推理时动态分配算力,提升大模型的逻辑推理能力

The Art of Scaling Test-Time Compute for Large Language Models

  • 通过大规模实验对比多种推理时算力扩展策略
  • 发现不同模型对算力的响应有明显差异,且性能随预算单调提升
  • 给出根据问题难易和模型类型选择最优策略的实用指南

测试时扩展(TTS)——即推理过程中动态分配算力——是提升大语言模型推理能力的有前景方向。然而,缺乏在相同条件下对知名TTS策略的系统比较,且模型类型与问题难度对性能的影响尚不明确。为此,我们开展了首个大规模TTS研究,覆盖超过三十亿个标记的生成内容,使用八种开源大模型(参数量7B至235B),在四个推理数据集上进行测试。观察到三个一致趋势:(1)没有单一TTS策略始终占优;(2)推理模型在不同问题难度和推理路径长度下表现出显著不同的轨迹质量,可划分为短视野与长视野两类;(3)对于特定模型类型,最优TTS性能随算力预算单调上升。基于这些发现,我们提供了一套实用策略选择方法,综合考虑问题难度、模型类型与算力预算,为有效推理时扩展提供指导。

原文摘要 · Abstract (English)

Test-time scaling (TTS) -- the dynamic allocation of compute during inference -- is a promising direction for improving reasoning in large language models (LLMs). However, a systematic comparison of well-known TTS strategies under identical conditions is missing, and the influence of model type and problem difficulty on performance remains unclear. To address these gaps, we conduct the first large-scale study of TTS, spanning over thirty billion tokens generated using eight open-source LLMs (7B to 235B parameters), across four reasoning datasets. We observe three consistent trends: (1) no single TTS strategy universally dominates; (2) reasoning models exhibit distinct trace-quality patterns across problem difficulty and trace length, forming short-horizon and long-horizon categories; and (3) for a given model type, the optimal TTS performance scales monotonically with compute budget. Based on these insights, we provide a practical recipe for selecting the best TTS strategy, considering problem difficulty, model type, and compute budget, providing a practical guide to effective inference-time scaling.

大模型推理算力扩展逻辑推理TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。