发现数学证明中的组合学习能力是攻克难题的关键,但非唯一决定因素。
On Compositional Learning Behaviours in Formal Mathematics
- 设计新评测框架S2B-LM,剥离数值干扰并用思维链激发真实组合能力。
- 实验显示抑制组合能力后解题率从32.3%骤降至2.9%,证明其必要性。
- 适合对自动定理证明、符号推理和模型可解释性感兴趣的读者。
能解决形式数学难题的自演化科学智能体需要具备组合学习行为(CLBs)——即在上下文中理解并重组新颖符号结构的能力,而不仅是复用已学原子。我们提出S2B-LM,一种改进的符号行为基准,消除数值处理的混淆因素,并引入链式思维支架以激发而非仅探测潜在的CLB能力。在十款Lean~4定理证明器上交叉评估其在S2B-LM中的CLB表现与miniF2F全证明性能,发现相关与因果证据:第一,通过四象限检验的必要条件分析得p=0.004,排除模型规模为混杂因子;第二,使用对比激活添加法从DeepSeek-Prover-V2-7B中提取一个编码CLB的激活方向,在AIME子集上用于miniF2F全证明生成时,抑制该方向导致求解率从32.3%跌至2.9%,且不损失连贯性,而抑制同量级随机方向则维持在31.9%。结果表明,CLB能力对攻克形式数学验证的长尾难题是必要但不充分的。
原文摘要 · Abstract (English)
Self-evolving scientific agents capable of conquering the hard tail of formal mathematics require Compositional Learning Behaviours (CLBs) -- the capacity to ground and recombine novel symbolic structures in context, beyond mere recombination of prelearned atoms. We propose S2B-LM, an adaptation of the CLB-evaluating Symbolic Behaviour Benchmark that removes numerical processing as a confound and adds chain-of-thought scaffolding to elicit rather than merely probe latent CLB competency. Cross-evaluating ten Lean~4 theorem provers on CLB competency in S2B-LM and miniF2F whole-proof performance, we find correlational and causal evidence of our claim: First, a necessary-condition analysis via quadrant test yields $p=0.004$, with model scale being ruled out as a confound. Second, extracting a CLB-encoding activation direction from DeepSeek-Prover-V2-7B using S2B-LM traces via Contrastive Activation Addition and applying it during miniF2F whole-proof generation on the AIME subset, CLB suppression collapses solve rate from $32.3\%$ to $2.9\%$, without loss of coherence, while suppressing a random activation direction of equal magnitude leaves it at $31.9\%$. Together, these results show that CLB competency is necessary but not sufficient for the hard tail of formal mathematical verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。