arXiv:2605.07600cs.LGcs.AI2026-05

用大模型自身模拟因果干预,精准识别解题关键概念。

Mathematical Reasoning via Intervention-Based Time-Series Causal Discovery Using LLMs as Concept Mastery Simulators

  • 让模型假设某概念已掌握,观察答案变化以判断其因果作用。
  • 实测显示正确解题问题的因果效应高出6.1倍,预测能力显著。
  • 适合研究模型知识激活机制或提升数学推理性能的研究者。

现有提升大模型数学推理的方法,如基于MCTS的测试时搜索或因果图引导的知识注入,无法区分真正起因果作用的概念,因观测关联可能受问题难度等混淆因素影响。本文提出CIKA(因果干预知识激活)框架,利用大模型自身作为干预模拟器:通过提示将某概念设为‘已掌握’,根据答案正确率的变化估算因果效应。该量定义为干预能力探测器(ICP),用于诊断模型是否真正能使用某概念——而非仅具备知识。由于干预外生设定概念状态,独立于问题难度,ICP可分离观测方法无法消除的混杂。在67个筛选问题上,排名第一概念的ICP值(+0.219)显著高于负向对照(+0.039;配对t检验,p < 10⁻⁶,Cohen's d = 0.86),验证探测器能有效区分因果相关与无关概念。对601个Omni-MATH问题分析进一步表明,解出问题的平均处理效应(ATE)为0.338,未解出问题仅为0.055,相差6.1倍,证实ICP可预测解题成功。使用参数量7B且权重冻结的模型,CIKA在无污染的Omni-MATH-Rule基准上达69.7%,整体达64.0%,优于o1-mini的60.5%;在GSM8K上达97.2%,在AIME 2024–2026上达46–50%,在MathArena上达46.2%。在基础模型失败的问题中,因果知识激活组件贡献了33.8%的正确答案,证明模型已具备但未激活所需知识。

原文摘要 · Abstract (English)

Recent methods for improving LLM mathematical reasoning, whether through MCTS-based test-time search or causal graph-guided knowledge injection, cannot identify which concepts causally contribute to a correct answer, as the observed association may be spurious, driven by confounders such as problem difficulty. We propose CIKA (Causal Intervention for Knowledge Activation), a framework that uses the LLM itself as an interventional simulator: a prompt sets the concept state to ``mastered'' and the correctness change estimates the causal effect. We formalize this quantity as an Interventional Capability Probe (ICP), which diagnoses whether the LLM can use a given concept -- distinct from merely possessing knowledge. Because the intervention exogenously sets the concept state independently of problem difficulty, ICP separates confounding that observational methods cannot. On 67 screened problems, the ICP of the top-ranked concept (+0.219) is significantly larger than that of the negative control (+0.039; paired $t$-test, $p < 10^{-6}$, Cohen's $d = 0.86$), confirming that the probe discriminates causally relevant concepts from irrelevant ones. Analysis of 601 Omni-MATH problems further shows that solved problems have 6.1$\times$ higher ATE than unsolved ones (0.338 vs. 0.055), confirming that ICP is predictive of problem-solving success. With a 7B-parameter LLM whose weights are entirely frozen, CIKA achieves 69.7\% on the contamination-free Omni-MATH-Rule benchmark and 64.0\% overall, compared to 60.5\% for o1-mini, and 97.2\% on GSM8K, 46--50\% on AIME 2024--2026, and 46.2\% on MathArena. The Causal Knowledge Activation component contributes 33.8\% of correct answers on problems where the base model alone fails, demonstrating that the LLM already possessed but had not activated the requisite knowledge.

数学推理因果发现知识激活

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。