通过推理时扩展提升大模型复杂任务能力,发现效果因任务而异且有上限。
Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
- 用多次调用与反馈机制模拟推理扩展,评估不同模型表现
- 高复杂度任务中扩展效果减弱,单纯增加token不保证准确率提升
- 完美验证器能显著提效,说明未来改进空间仍大
推理时扩展可增强大语言模型在需逐步推理的复杂任务中的能力。尽管延长思维链对数学任务有效,其在其他任务上的影响尚不明确。本文研究了九个前沿模型和八个挑战性任务(包括数学与理工科推理、日历规划、NP难问题、导航与空间推理)中扩展方法的优劣。通过重复调用模型(独立或带反馈)评估常规模型(如GPT-4o)与专为推理扩展优化的模型(如o1)的表现,近似各模型的性能下界与上界,以及未来提升潜力。实证分析表明,推理扩展的优势随任务复杂度增加而减弱,且仅增加生成长度并不必然提高准确率。多轮独立实验显示,部分任务中常规模型结合完美验证器可达先进模型平均表现;但在另一些任务中,即便在极高扩展规模下仍存在显著差距。令人鼓舞的是,所有模型在引入完美验证器或强反馈后均实现显著提升,表明未来仍有巨大优化空间。
原文摘要 · Abstract (English)
Inference-time scaling can enhance the reasoning capabilities of large language models (LLMs) on complex problems that benefit from step-by-step problem solving. Although lengthening generated scratchpads has proven effective for mathematical tasks, the broader impact of this approach on other tasks remains less clear. In this work, we investigate the benefits and limitations of scaling methods across nine state-of-the-art models and eight challenging tasks, including math and STEM reasoning, calendar planning, NP-hard problems, navigation, and spatial reasoning. We compare conventional models (e.g., GPT-4o) with models fine-tuned for inference-time scaling (e.g., o1) through evaluation protocols that involve repeated model calls, either independently or sequentially with feedback. These evaluations approximate lower and upper performance bounds and potential for future performance improvements for each model, whether through enhanced training or multi-model inference systems. Our extensive empirical analysis reveals that the advantages of inference-time scaling vary across tasks and diminish as problem complexity increases. In addition, simply using more tokens does not necessarily translate to higher accuracy in these challenging regimes. Results from multiple independent runs with conventional models using perfect verifiers show that, for some tasks, these models can achieve performance close to the average performance of today's most advanced reasoning models. However, for other tasks, a significant performance gap remains, even in very high scaling regimes. Encouragingly, all models demonstrate significant gains when inference is further scaled with perfect verifiers or strong feedback, suggesting ample potential for future improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。