arXiv:2606.13473cs.LGcs.AI2026-06被引 2

用群体测试时扩展让模型在数学竞赛中超越人类金牌水平

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

论文配图:MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
图 1 · 摘自论文原文
  • 训练生成、验证、修复三能力,用低误报验证器确保可靠性
  • 测试时通过群体候选证明筛选,获IMO 2025 35/42、USAMO 2026 36/42
  • 适合研究数学证明自动化与强化学习推理的学者和竞赛选手

我们提出 MaxProof,一种面向 MiniMax-M3 系列竞赛级数学证明的群体级测试时扩展框架。M3 首先通过设计精良的生成式验证器(低误报率)训练三种证明导向能力:证明生成、证明验证与批判条件下的证明修复。这些能力整合为单一发布的 M3 模型。测试时,MaxProof 将该模型视为生成器、验证器、优化器与评分器,对候选证明群体进行搜索,并通过锦标赛选择返回最终证明。借助 MaxProof 测试时扩展,M3 模型在 IMO 2025 上取得 35/42 分,在 USAMO 2026 上取得 36/42 分,均超过人类金牌标准。

原文摘要 · Abstract (English)

We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.

数学证明强化学习测试时扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。