arXiv:2503.24235cs.CLcs.AI2025-03综述被引 183

系统梳理大模型测试时扩增技术,讲清做什么、怎么做、在哪用、效果如何。

A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

  • 按四维度(内容、方法、位置、效果)构建统一分析框架
  • 揭示测试时扩增可显著提升数学、编程等任务性能
  • 适合研究者和工程师快速掌握该领域全貌

随着预训练阶段算力扩展热情减退,测试时扩增(Test-Time Scaling, TTS)——又称“测试时计算”——成为研究热点。近期研究表明,TTS能进一步激发大语言模型(LLMs)的问题求解能力,在数学、编程等专项推理任务以及开放问答等通用任务中均取得显著突破。然而,该领域研究迅速增长,仍缺乏系统性综述。为此,本文提出一个四维统一框架:何物可扩、如何扩、何处扩、扩得如何。基于此,我们全面回顾方法、应用场景与评估维度,梳理各类技术在整体架构中的功能角色。通过分析,提炼出TTS的主要发展脉络,并提供实用部署指南。同时指出若干开放挑战与未来方向,包括进一步扩增、明晰技术本质、拓展任务泛化性及可解释性。相关资源已开源至 https://github.com/testtimescaling/testtimescaling.github.io/

原文摘要 · Abstract (English)

As enthusiasm for scaling computation (data and parameters) in the pretraining era gradually diminished, test-time scaling (TTS), also referred to as ``test-time computing'' has emerged as a prominent research focus. Recent studies demonstrate that TTS can further elicit the problem-solving capabilities of large language models (LLMs), enabling significant breakthroughs not only in specialized reasoning tasks, such as mathematics and coding, but also in general tasks like open-ended Q&A. However, despite the explosion of recent efforts in this area, there remains an urgent need for a comprehensive survey offering a systemic understanding. To fill this gap, we propose a unified, multidimensional framework structured along four core dimensions of TTS research: what to scale, how to scale, where to scale, and how well to scale. Building upon this taxonomy, we conduct an extensive review of methods, application scenarios, and assessment aspects, and present an organized decomposition that highlights the unique functional roles of individual techniques within the broader TTS landscape. From this analysis, we distill the major developmental trajectories of TTS to date and offer hands-on guidelines for practical deployment. Furthermore, we identify several open challenges and offer insights into promising future directions, including further scaling, clarifying the functional essence of techniques, generalizing to more tasks, and more attributions. Our repository is available on https://github.com/testtimescaling/testtimescaling.github.io/

大模型测试时扩增综述推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。