将数学推理拆解为原子能力,揭示大模型真实认知机制
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- 把数学能力拆解为领域与逻辑双维度的原子单元
- 发现不同原子能力间存在显著影响与交互关系
- 适合研究模型认知、提升推理可解释性的学者
大语言模型在数学推理上表现优异,但当前主流方法依赖大规模数据与长思考链,引发其是否真正掌握数学概念而非仅记忆训练数据的质疑。人类则擅长将复杂问题分解为基本原子能力。受此启发,本文提出评估数学原子能力的新范式,将能力划分为四大领域(代数、几何、分析、拓扑)和三类逻辑层次(概念理解、多步形式化推理、反例驱动的逆向推理)。针对每个原子能力单元设计训练与评估数据集,并在先进模型上开展实验,探究不同能力间的相互影响。结果揭示模型在各类原子能力上的表现差异及交互模式,强调将数学智能解耦为原子组件的重要性,为理解模型认知、发展更高效、可迁移、具认知基础的'原子思维'训练范式提供新思路。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated outstanding performance in mathematical reasoning capabilities. However, we argue that current large-scale reasoning models primarily rely on scaling up training datasets with diverse mathematical problems and long thinking chains, which raises questions about whether LLMs genuinely acquire mathematical concepts and reasoning principles or merely remember the training data. In contrast, humans tend to break down complex problems into multiple fundamental atomic capabilities. Inspired by this, we propose a new paradigm for evaluating mathematical atomic capabilities. Our work categorizes atomic abilities into two dimensions: (1) field-specific abilities across four major mathematical fields, algebra, geometry, analysis, and topology, and (2) logical abilities at different levels, including conceptual understanding, forward multi-step reasoning with formal math language, and counterexample-driven backward reasoning. We propose corresponding training and evaluation datasets for each atomic capability unit, and conduct extensive experiments about how different atomic capabilities influence others, to explore the strategies to elicit the required specific atomic capability. Evaluation and experimental results on advanced models show many interesting discoveries and inspirations about the different performances of models on various atomic capabilities and the interactions between atomic capabilities. Our findings highlight the importance of decoupling mathematical intelligence into atomic components, providing new insights into model cognition and guiding the development of training strategies toward a more efficient, transferable, and cognitively grounded paradigm of "atomic thinking".
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。