通过构建因果推理图,揭示大模型如何隐式理解数学推理结构。
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
- 用因果图建模语言模型推理过程中的细粒度依赖关系。
- 实验证明模型输出节点对最终答案有因果贡献,且重视图中路径。
- 适合研究大模型推理机制的学者,可支持可控干预实验。
链式思维(CoT)已被证明能提升大语言模型在多种推理任务上的表现,但其增益机制尚无定论。为此,我们提出因果链式思维图(CCGraphs),一种从推理轨迹自动提取的有向无环图,用于刻画语言模型输出中的细粒度因果依赖关系。我们构建了包含1671道来自MATH500、GSM8K和AIME的数学推理题及其对应CCGraphs的数据集——KisMATH。基于15个开源权重大模型的实证分析显示:(i) CCGraph中的推理节点是最终答案的因果贡献者,这构成了推理的核心;(ii) 大模型倾向于强调CCGraph所捕捉的推理路径,表明模型内部实现了与我们图结构相似的表征。KisMATH支持受控的、图对齐的干预,为深入探究CoT在大模型推理中的作用开辟了新路径。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) traces have been shown to improve performance of large language models on a plethora of reasoning tasks, yet there is no consensus on the mechanism by which this boost is achieved. To shed more light on this, we introduce Causal CoT Graphs (CCGraphs), which are directed acyclic graphs automatically extracted from reasoning traces that model fine-grained causal dependencies in language-model outputs. A collection of 1671 mathematical reasoning problems from MATH500, GSM8K, and AIME, together with their associated CCGraphs, has been compiled into our dataset -- KisMATH. Our detailed empirical analysis with 15 open-weight LLMs shows that (i) reasoning nodes in the CCGraphs are causal contributors to the final answer, which we argue is constitutive of reasoning; and (ii) LLMs emphasize the reasoning paths captured by the CCGraphs, indicating that the models internally realize structures similar to our graphs. KisMATH enables controlled, graph-aligned interventions and opens avenues for further investigation into the role of CoT in LLM reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。