用结构化图谱分析大模型推理路径,揭示传统指标掩盖的思维差异。
Reasoning Structure of Large Language Models

- 将模型推理过程转为可验证的命题依赖图
- 提出衡量逻辑流集中度的效率指标
- 能区分看似相同但结构不同的推理方式
大型推理模型常以最终答案准确率或词元数量评估,但相同评分可能隐藏本质不同的推理结构。为此,我们构建了一个可扩展的逻辑谜题基准和一套将非结构化推理轨迹转换为可验证的命题与依赖关系图的流程。该方法使推理成为可量化分析的结构化对象。基于此,我们定义了推理效率指标,衡量模型逻辑流的集中程度。对开源推理模型的分析表明,结构测量能分离出被词元数和准确率混淆的行为,为诊断失败模式和比较推理随题目难度的变化提供实用工具。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics can hide fundamentally different reasoning structures. To address this limitation, we introduce a scalable LRM benchmark of logic puzzles and a pipeline that converts unstructured traces into verifiable reasoning graphs of claims and dependencies. This turns reasoning into a structured, measurable object whose topology can be quantitatively analyzed. Building on this, we define a reasoning efficiency metric that quantifies how concentrated the model's logical flow is. Our analysis on open-source reasoning models shows that structural measurements separate behaviors that token count and accuracy conflate, providing a practical tool for diagnosing failure modes and comparing how reasoning scales with puzzle difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。