arXiv:2506.21609cs.CLcs.AI2025-06被引 1

对比四款顶级推理模型,揭示其思考与输出的差异。

From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models

  • 用关键词统计和大模型评判法分析思维过程与输出关联
  • 发现模型在推理深度、中间步骤依赖度上存在显著差异
  • 适合关注模型设计与评估的开发者及研究者参考

近期大型语言模型(LLMs)在复杂推理能力上取得显著进展,但现有研究普遍缺乏对其推理过程与输出的系统性对比,尤其忽视了自我反思模式(即“顿悟时刻”)以及跨领域关联性。本文提出一种新框架,通过关键词统计与大模型作为裁判(LLM-as-a-judge)范式,分析GPT-o1、DeepSeek-R1、Kimi-k1.5和Grok-3四款前沿推理模型的推理特征。研究基于涵盖逻辑推理、因果推断与多步求解的真实场景问题数据集,引入多项指标评估推理连贯性与输出准确性。结果揭示各模型在探索与利用之间的平衡、问题处理策略及结论生成方式上的多样化模式。定量与定性比较表明,不同模型在推理深度、中间步骤依赖程度,以及思维过程与输出模式相似性方面存在明显差异。本研究为计算效率与推理鲁棒性间的权衡提供洞见,并为实际应用中的模型设计与评估提供实践建议。项目已公开于:https://github.com/ChangWenhan/FromThinking2Output

原文摘要 · Abstract (English)

Recently, there have been notable advancements in large language models (LLMs), demonstrating their growing abilities in complex reasoning. However, existing research largely overlooks a thorough and systematic comparison of these models' reasoning processes and outputs, particularly regarding their self-reflection pattern (also termed "Aha moment") and the interconnections across diverse domains. This paper proposes a novel framework for analyzing the reasoning characteristics of four cutting-edge large reasoning models (GPT-o1, DeepSeek-R1, Kimi-k1.5, and Grok-3) using keywords statistic and LLM-as-a-judge paradigm. Our approach connects their internal thinking processes with their final outputs. A diverse dataset consists of real-world scenario-based questions covering logical deduction, causal inference, and multi-step problem-solving. Additionally, a set of metrics is put forward to assess both the coherence of reasoning and the accuracy of the outputs. The research results uncover various patterns of how these models balance exploration and exploitation, deal with problems, and reach conclusions during the reasoning process. Through quantitative and qualitative comparisons, disparities among these models are identified in aspects such as the depth of reasoning, the reliance on intermediate steps, and the degree of similarity between their thinking processes and output patterns and those of GPT-o1. This work offers valuable insights into the trade-off between computational efficiency and reasoning robustness and provides practical recommendations for enhancing model design and evaluation in practical applications. We publicly release our project at: https://github.com/ChangWenhan/FromThinking2Output

推理模型思维链模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。