arXiv:2601.16172cs.AI2026-01中稿 · ICML

强化学习训练的证明器推理时会陷入模式崩溃,通过引入结构多样性可显著提升效果。

Inference-Time Diversity in RL-Trained Lean Theorem Provers: A Diagnostic Study

  • 用固定策略骨架打破推理时的模式崩溃,恢复证明能力。
  • 在 $k=16$ 时相对提升 45%,平均多解 12.3 个定理($n=3$ 种随机种子)。
  • 结构多样性低成本有效,适合改进强化学习训练的数学证明器。

强化学习训练的 Lean 证明器在推理时出现模式崩溃:在 miniF2F 测试上,将独立同分布采样预算从 $k{=}32$ 提升至 $k{=}64$ 未增加任何新解(均为 42/244)。采用固定 15 个策略骨架后,在 $k{=}16$ 时实现 +45% 相对提升(均值 $Δ= +12.3 \pm 4.2$ 个定理,$n{=}3$ 种随机种子,每种种子符号一致)。控制性消融实验排除了提示多样性干扰:策略骨架有效,改写语句表现与基线持平,无关的 Lean 注释反而降低性能。留一法形式化难度分层揭示三种扰动间存在结构-内容梯度。该现象为 RL 特有:V1.5-Base 即使干预也解不出任何定理,表明是强化学习阶段赋予了初始证明能力,但后续发生坍塌;扩展至两个额外 7B 模型,强化学习训练的 DeepSeek-Prover-V2-7B 在 $k{=}64$ 时比所有独立同分布基线多解决 3 个前沿定理,而监督微调训练的 Goedel-Prover 性能下降 $-10.0 \pm 4.4$($n{=}3$,符号一致)。推理时的结构多样性是强化学习训练证明器的一种廉价、互补的优化方向,与模型规模或训练算力扩展正交。

原文摘要 · Abstract (English)

RL-trained Lean theorem provers mode-collapse at inference time: on miniF2F-test with DeepSeek-Prover-V1.5-RL, doubling the i.i.d.\ sampling budget from $k{=}32$ to $k{=}64$ produces zero additional solved theorems (42/244 in both cases). A fixed schedule of 15 tactic skeletons breaks this plateau and recovers a $+45%$ relative improvement at $k{=}16$ (mean $Δ= +12.3 \pm 4.2$ theorems across $n{=}3$ seeds, sign preserved in every seed). A controlled diversity ablation rules out the prompt-diversity confound: tactic skeletons help, paraphrases match the baseline, and irrelevant Lean comments actively degrade. A leave-one-out formalization-difficulty stratification reveals a structural-content gradient across the three perturbations. The phenomenon is RL-specific: V1.5-Base proves zero theorems regardless of intervention, identifying RL as the stage that creates the proof capability which subsequently collapses; extending to two additional 7B Lean provers, RL-trained DeepSeek-Prover-V2-7B contributes $+3$ frontier solves no i.i.d. baseline can reach despite a flat aggregate, while SFT-trained Goedel-Prover does not ( $-10.0$ $\pm 4.4$ theorems, $n{=}3$, sign preserved every seed). Inference-time structural diversity is a cheap, complementary axis for RL-trained provers, orthogonal to scaling model size or training compute.

强化学习形式化证明推理多样性数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。