arXiv:2608.23205cs.AIcs.CL2026-08

用布鲁姆分类法分析大模型推理步骤,揭示思维层级与正确性的关系。

Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy

论文配图:Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy
图 1 · 摘自论文原文
  • 基于布鲁姆分类法自动标注推理步骤的思维层级
  • 发现不同模型在任务中表现出相似但有差异的思维模式
  • 推理中的思维类型与结果正确性相关,可用于提升模型表现

大型推理模型(LRMs)革新了大语言模型的推理能力,随着推理过程记录的公开,研究模型行为不仅限于表层,还可深入到单个推理步骤。然而,理解推理中所采用的思维方式——这对揭示模型推理模式并实现实际应用至关重要——仍处于探索阶段。为此,本文提出一个基于布鲁姆分类法的自动标注框架,将思维分为六个认知层级(如记忆、应用、评估等)。通过该框架,在多个模型和数据集上开展大规模分析,揭示了模型与任务间思维模式的异同。进一步证明,从推理轨迹中提取的思维类型信息与推理正确性存在关联,为改进推理能力提供新路径。研究建立了一个细粒度的分析框架,为提升推理质量提供了可操作的洞见。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public availability of reasoning traces creates valuable opportunities to study model behavior not only at the surface level but also at the granularity of individual reasoning steps. However, understanding the types of thinking employed during reasoning - which offers critical insights into models' reasoning patterns and enables actionable applications - remains underexplored. To address this gap, we introduce a framework for automatic annotation of reasoning steps through the lens of Bloom's Taxonomy, which classifies thinking into six cognitive levels, such as Remembering, Applying and Evaluating. Using this framework, we perform a large-scale analysis across models and datasets, revealing both similarities and differences in thinking patterns across models and tasks. Moreover, we demonstrate that thinking-type information derived from reasoning traces correlates with correctness, paving the way for improved reasoning. Our findings establish a fine-grained framework for analyzing thinking patterns in LRMs and provide actionable insights for enhancing reasoning quality.

推理分析认知层级模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。