测试推理阶段计算对大模型思维与规划能力的提升效果
Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
- 构建了涵盖11项任务的综合评测基准Sys2Bench
- 发现单纯增加推理计算量无法统一提升各类任务表现
- 揭示计算开销与性能之间的权衡,适合研究推理优化者
我们考察大语言模型(LLMs)在解决复杂任务时的推理与规划能力。近期推理阶段技术的进步表明,通过在推理过程中探索中间步骤,可在不进行额外训练的情况下增强LLM的推理能力。值得注意的是,OpenAI的o1模型通过多步推理与验证实现了优异表现。本文研究了推理阶段技术的规模扩展如何改善推理与规划,重点在于理解计算成本与性能之间的权衡。为此,我们构建了一个综合性基准Sys2Bench,对现有推理阶段技术在五个类别共十一项多样化任务上进行了广泛实验,涵盖算术推理、逻辑推理、常识推理、算法推理和规划。结果表明,单纯扩大推理阶段计算存在局限性,没有单一推理技术能在所有任务上持续表现优异。
原文摘要 · Abstract (English)
We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential to enhance LLM reasoning without additional training by exploring intermediate steps during inference. Notably, OpenAI's o1 model shows promising performance through its novel use of multi-step reasoning and verification. Here, we explore how scaling inference-time techniques can improve reasoning and planning, focusing on understanding the tradeoff between computational cost and performance. To this end, we construct a comprehensive benchmark, known as Sys2Bench, and perform extensive experiments evaluating existing inference-time techniques on eleven diverse tasks across five categories, including arithmetic reasoning, logical reasoning, common sense reasoning, algorithmic reasoning, and planning. Our findings indicate that simply scaling inference-time computation has limitations, as no single inference-time technique consistently performs well across all reasoning and planning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。