arXiv:2605.19762cs.AIcs.CL2026-05中稿 · ICML被引 1

代码不提升通用推理,但结构化混合数据能显著增强数学推理。

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code

论文配图:What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
图 1 · 摘自论文原文
  • 用混合代码-文本和数学-文本数据提升推理能力,而非单纯依赖可执行代码。
  • 在固定数学数据预算下,增加结构化数学样本密度,数学推理性能提升显著。
  • 适合需要精准优化推理能力的研究者,尤其关注数学与编程的平衡问题。

代码已成为现代基础语言模型训练的标准成分,但其在编程之外的作用仍不明确。我们通过在10T token语料库上的受控预训练实验重新审视了这一问题,发现三方面结果:第一,当代码仅限于独立可执行程序且代码-自然语言数据被控制时,代码显著提升编程能力,但并非通用推理增强手段,反而与知识密集型任务竞争,尤其是复杂数学推理;第二,以往归因于代码的推理提升,更应归因于跨域结构化推理信号,如代码-文本和数学-文本混合数据;第三,在固定数学数据预算下,提高结构化数学样本密度可在显著提升困难数学推理表现的同时,基本保持编程性能,表明认知支架是缓解跨域权衡的有效策略。最后,路由分析显示,数据组合效应反映在专家激活模式中,提供了领域间竞争与协同作用的机制层面证据。研究明确了哪些数据特征能跨能力维度迁移,并指向更精确的数据中心优化策略。

原文摘要 · Abstract (English)

Code has become a standard component of modern foundation language model (LM) training, yet its role beyond programming remains unclear. We revisit the claim that code improves reasoning through controlled pretraining experiments on a 10T-token corpus with fine-grained domain separation. Our findings are threefold. First, when code is restricted to standalone executable programs and Code-NL data are controlled for, code substantially improves programming ability but does not act as a general reasoning enhancer; instead, it competes with knowledge-intensive tasks, especially complex mathematical reasoning. Second, the reasoning gains often attributed to code are better explained by cross-domain structured reasoning traces, such as code-text and math-text mixtures, rather than by executable code alone. Third, increasing the density of structured math-domain samples within a fixed math budget yields substantial gains on difficult mathematical reasoning while largely preserving programming performance, suggesting that cognitive scaffolds offer a targeted way to mitigate cross-domain trade-offs. Finally, routing analyses show that data-composition effects are reflected in expert-activation patterns, providing mechanism-level evidence for competitive and synergistic interactions across domains. Our results clarify which data characteristics transfer across capability dimensions and point to more precise data-centric optimization strategies.

数学推理结构化数据模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。