arXiv:2505.13307cs.CLcs.AI2025-05TPAMI被引 6

提出可量化推理边界的框架,解决大模型复杂推理的评估与优化难题。

RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning

  • 定义推理边界(RB)为思维链性能上限,提出组合定律实现跨任务量化分析。
  • 在38个模型、13个任务上验证框架有效性,支持多模态感知等不可测能力评估。
  • 提供10种思维链策略对比,助力理解推理优化与退化机制,适合模型开发者参考。

思维链(CoT)推理在提升大语言模型处理复杂任务的能力方面已证明有效,但其实际应用仍面临两大挑战:一是缺乏对可测量推理能力边界的定量评估指标与优化指导;二是难以评估不可测量能力(如多模态感知)的边界。为此,本文提出推理边界框架++(RBF++)。针对第一类问题,定义推理边界(RB)为思维链性能的最大极限,并提出RB组合定律,支持跨任务的量化分析与可操作指导。针对第二类问题,引入恒定假设,将不可测的RB替换为场景特定常数;同时提出推理边界划分机制,将不可测边界拆分为两个子边界,实现对不可测领域知识与多模态感知能力的量化与优化。在涵盖38个模型和13个任务的大量实验中验证了框架在跨模态场景下的可行性。此外,评估了10种CoT策略,从互补视角揭示优化与退化规律,并扩展了用于衡量LLM推理边界的评估基准。代码与数据已开源。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models (LLMs) on complex tasks, spurring research into its underlying mechanisms. However, two primary challenges remain for real-world applications: (1) the lack of quantitative metrics and actionable guidelines for evaluating and optimizing measurable boundaries of CoT capability, and (2) the absence of methods to assess boundaries of unmeasurable CoT capability, such as multimodal perception. To address these gaps, we introduce the Reasoning Boundary Framework++ (RBF++). To tackle the first challenge, we define the reasoning boundary (RB) as the maximum limit of CoT performance. We also propose a combination law for RBs, enabling quantitative analysis and offering actionable guidance across various CoT tasks. For the second challenge, particularly in multimodal scenarios, we introduce a constant assumption, which replaces unmeasurable RBs with scenario-specific constants. Additionally, we propose the reasoning boundary division mechanism, which divides unmeasurable RBs into two sub-boundaries, facilitating the quantification and optimization of both unmeasurable domain knowledge and multimodal perception capabilities. Extensive experiments involving 38 models across 13 tasks validate the feasibility of our framework in cross-modal settings. Additionally, we evaluate 10 CoT strategies, offer insights into optimization and decay from two complementary perspectives, and expand evaluation benchmarks for measuring RBs in LLM reasoning. We hope this work advances the understanding of RBs and optimization strategies in LLMs. Code and data are available at https://github.com/LightChen233/reasoning-boundary.

推理边界思维链大模型评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。