用风险控制方法评估自然语言数学答案,提升可信度。
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning

- 设计COVCAL选择器,基于Lean诊断结果控制接受风险。
- 7B模型仅覆盖28%问题,证明可信度仅43%;专用模型达79%覆盖率。
- 在高覆盖率下可安全接受48%问题,准确率达98%,适合谨慎验证场景。
Lean被越来越多用于判断自然语言数学答案的正确性,但其信号具有局限性:许多答案无法形式化,且证明失败可能源于类型错误或缺少库事实,而非答案本身错误。在MATH-500数据集上,我们发现该信号具有显著的覆盖率依赖性——高覆盖率下正确率可达96%,低覆盖率时仅20%。此外,信号稀疏且常不忠实:7B自动形式化器仅对28%的问题生成形式化内容,人工审计显示其中约43%的证明是可信的。我们提出COVCAL,一种基于Lean追踪诊断的选择器,可在两种策略下(保守的邦弗朗尼界与更紧的先偏差后校准规则)为被接受的答案提供有限样本的风险保证,或选择不答。可行性取决于自动形式化覆盖率:使用7B形式化器时信号过稀疏,邦弗朗尼策略在所有20个自助样本中均选择不答;而采用专为证明优化的形式化器可达到79%覆盖率,在17/20个样本中实现可行,接受约48%的问题,且接受准确率高达0.98。由于自一致性本身已达91%准确率,本工作的贡献在于精确界定在何种形式化器和条件下,可基于部分形式化信号进行可控风险的信任。
原文摘要 · Abstract (English)
Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wrong answer. On MATH-500 we show this signal is (i) sharply coverage-dependent, that is the proof-winning answer is correct 96% of the time at high proved coverage but 20% at low, and (ii) sparse and often unfaithful: a 7B autoformalizer proves a class for only 28% of problems, and a manual audit finds only approximately 43% of those proofs faithful. We propose COVCAL, a selector over Lean-trace diagnostics that certifies a finite-sample selective-risk bound on accepted answers or abstains, under two regimes (a conservative Bonferroni bound and a tighter dev-then-cal rule). Feasibility depends on autoformalization coverage: with the 7B formalizer the signal is too sparse and Bonferroni abstains on all 20 bootstrap partitions, whereas a prover-specialized formalizer reaches 79% coverage and flips it to feasible on 17 of 20, accepting approximately 48% of problems at 0.98 accepted accuracy. Since self-consistency alone is already 91% accurate, our contribution is a precise account of when, and with which formalizer, a partial formal signal can be trusted under risk control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。