小模型也能精准验证数学证明,只需特定提示词。
Do We Need Frontier Models to Verify Mathematical Proofs?
- 用自研提示词优化小模型,突破其验证能力瓶颈
- 小模型准确率仅比前沿模型低10%,但一致性差25%
- 适合需要低成本、高可靠性的数学验证场景
训练、后训练和推理时方法的进步使前沿推理模型在数学竞赛中获得金牌,并解决复杂开放问题。信任这些模型的输出需对自然语言证明进行错误检查。大语言模型裁判正被广泛采用以满足日益增长的评估需求。尽管验证被认为比生成更容易,但可靠验证究竟需要什么模型能力?我们系统评估了四种开源和两种前沿大模型在人类评分的竞赛级问题自然语言证明数据集上的表现。重点关注两个指标:验证器准确率与自一致性(同一证明多次判断的一致率)。结果发现,小模型准确率仅落后前沿模型约10%,但自一致性差距高达25%。此外,所有模型的验证准确率均对提示词选择敏感。我们进一步证明,小模型实际上具备与前沿模型相当的数学验证能力,但难以通过通用提示词稳定激发。通过大模型引导的提示词搜索,我们设计出一组专用提示词,有效克服小模型的特定失败模式,使其准确率最高提升9.1%,自一致性提升15.9%。该效果在不同模型和数据集上均成立,使Qwen3.5-35B的表现达到与Gemini 3.1 Pro相当的水平。
原文摘要 · Abstract (English)
Advances in training, post-training, and inference-time methods have enabled frontier reasoning models to win gold medals in math competitions and settle challenging open problems. Gaining trust in the responses of these models requires that natural language proofs be checked for errors. LLM judges are increasingly being adopted to meet the growing demand for evaluating such proofs. While verification is considered easier than generation, what model capability does reliable verification actually require? We systematically evaluate four open-source and two frontier LLMs on datasets of human-graded natural language proofs of competition-level problems. We consider two key metrics: verifier accuracy and self-consistency (the rate of agreement across repeated judgments on the same proof). We observe that smaller open-source models are only up to ~10% behind frontier models in accuracy but they are up to ~25% more inconsistent. Furthermore, we see that verifier accuracy is sensitive to prompt choice across all models. We then demonstrate that the smaller models, in fact, do possess the mathematical capabilities to verify proofs at the level of frontier models, but they struggle to reliably elicit these capabilities with general judging prompts. Through an LLM-guided prompt search, we synthesize an ensemble of specialized prompts that overcome the specific failure modes of smaller models, boosting their performance by up to 9.1% in accuracy and 15.9% in self-consistency. These gains are realized across models and datasets, allowing models like Qwen3.5-35B to perform on par with frontier models such as Gemini 3.1 Pro for proof verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。