arXiv:2410.05695cs.CL2024-10NeurIPS被引 103

提出推理边界框架,量化大模型的思维链能力并指导优化。

Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought

  • 定义推理边界(RB)量化思维链上限,建立组合定律实现可计算评估。
  • 在27个模型、5项任务上验证框架有效性,解释10种思维链策略原理。
  • 提供从路径优化到能力提升的双重视角,适合模型开发者与研究者。

思维链(Chain-of-Thought, CoT)推理已成为提升大语言模型(LLMs)在复杂推理任务中表现的有力方法。然而,现有研究面临两大挑战:缺乏对CoT能力的定量评估指标,以及缺少优化指导。为此,本文提出一种新型推理边界框架(RBF),以解决上述问题。为实现量化评估,我们首先定义推理边界(RB)来衡量CoT的理论上限,并建立其组合定律,使该方法可应用于多种真实场景下的CoT任务。针对优化难题,我们划分出三类RB,并基于组合定律设计促进策略与推理路径优化方案。通过在27个模型和5个任务上的广泛实验,验证了该框架的存在性与合理性。同时,框架成功解释了10种典型CoT策略的有效性,并从两个维度指导优化。我们希望本工作能为理解大模型推理的边界与优化路径提供全面支持。代码与数据已开源于https://github.com/LightChen233/reasoning-boundary。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) reasoning has emerged as a promising approach for enhancing the performance of large language models (LLMs) on complex reasoning tasks. Recently, a series of studies attempt to explain the mechanisms underlying CoT, aiming to deepen the understanding of its efficacy. Nevertheless, the existing research faces two major challenges: (1) a lack of quantitative metrics to assess CoT capabilities and (2) a dearth of guidance on optimizing CoT performance. Motivated by this, in this work, we introduce a novel reasoning boundary framework (RBF) to address these challenges. To solve the lack of quantification, we first define a reasoning boundary (RB) to quantify the upper-bound of CoT and establish a combination law for RB, enabling a practical quantitative approach applicable to various real-world CoT tasks. To address the lack of optimization, we propose three categories of RBs. We further optimize these categories with combination laws focused on RB promotion and reasoning path optimization for CoT improvement. Through extensive experiments on 27 models and 5 tasks, the study validates the existence and rationality of the proposed framework. Furthermore, it explains the effectiveness of 10 CoT strategies and guides optimization from two perspectives. We hope this work can provide a comprehensive understanding of the boundaries and optimization strategies for reasoning in LLMs. Our code and data are available at https://github.com/LightChen233/reasoning-boundary.

推理边界思维链大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。