用时间锚定的证明路径新颖性评分,让形式化数学中的创新可量化。
PriorProof: A Point-in-Time Measure of Technique Novelty for Formal Proofs
- 基于早期数学库快照构建先验模型,自动计算证明路径的意外程度。
- 在76组对比中,评分一致性达69.7%,对关键配对的判断准确率超90%。
- 适合关注形式化数学创新性评估的研究者,可作为专家判断的辅助信号。
数学家常区分有解释力、简化或引入非标准路径的证明,但这类判断难以操作化。本文聚焦更窄的构念:形式数学中相对于时间的证明路径非标准性。对于一个Lean定理,PriorProof提取其展开证明项的依赖足迹,并在仅基于Mathlib前期季度快照构建的检索条件化、分层平滑先验下,对足迹的加权意外度进行评分。该方法无需手工构建技术本体或人工标注:语句检索通过从证明衍生的对比对中学习,评分对象则直接从证明项中机械读取。在盲态拓扑研究中,100个演示缩减为76个不同基础配对:12个标准对比重复三次用于一致性筛查,64个分层配对独立。与多数保留领域评审员相比,PriorProof在76对中达成53对一致(69.7%,95%置信区间58.7-78.9%),包括12对标准对比中的11对(91.7%,64.6-98.5%)和64对分层对比中的42对(65.6%,53.4-76.1%)。得分差距四分位数在重复合并后呈非单调分布;最小差距箱为12/19(63.2%,41.0-80.9%),最大差距箱为16/19(84.2%,62.4-94.5%),支持端点校准倾向而非阶梯式分辨。最佳语言模型条件在76对中达成60对一致(78.9%,68.5-86.6%);在成对结果上,PriorProof单独正确8对,模型单独正确15对(双侧精确McNemar检验p=0.210),表明当前样本量下差异未确立。因此,本文提出PriorProof并非替代专家或模型判断,而是一种可分解、时间锚定的信号,其得分差距可提供可解释的可靠性指标。
原文摘要 · Abstract (English)
Mathematicians distinguish proofs that explain, simplify, or introduce a nonstandard route, but these judgments are difficult to operationalize. We study a deliberately narrower construct: time-relative proof-route nonstandardness in formal mathematics. For a Lean theorem, PriorProof extracts the dependency footprint of its elaborated proof term and scores the weighted surprisal of that footprint under a retrieval-conditioned, hierarchically smoothed prior built only from an earlier quarterly snapshot of Mathlib. The method requires no hand-built technique ontology and no human labels: statement retrieval is learned from proof-derived contrastive pairs, while the scored object is read mechanically from proof terms. In a blinded topology study, 100 presentations collapse to 76 distinct underlying pairs: 12 canonical contrasts shown three times for consistency screening and 64 distinct stratified pairs. Against the majority of three retained domain raters, PriorProof agrees on 53/76 pairs (69.7%, Wilson 95% CI 58.7-78.9%), including 11/12 canonical pairs (91.7%, 64.6-98.5%) and 42/64 stratified pairs (65.6%, 53.4-76.1%). Score-gap quartiles are nonmonotone after repeat collapse; the endpoints are 12/19 (63.2%, 41.0-80.9%) in the smallest-gap bin and 16/19 (84.2%, 62.4-94.5%) in the largest, supporting an endpoint-calibration tendency rather than a resolved staircase. The best language-model condition agrees on 60/76 pairs (78.9%, 68.5-86.6%); on paired outcomes, PriorProof alone is correct on 8 pairs and the model alone on 15 (exact two-sided McNemar p = 0.210), so the difference is not established at this sample size. We therefore present PriorProof not as a replacement for expert or model judgment, but as a decomposable, time-anchored signal whose score gap provides an interpretable reliability indicator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。