用通用验证优化流程,让AI在奥数赛中准确率提升至85.7%。
Winning Gold at IMO 2025 with a Model-Agnostic Verification-and-Refinement Pipeline
- 设计不依赖模型的验证与迭代流程,自动纠错并优化解题路径。
- 在IMO 2025上,三款主流模型经该流程后正确解决5道题(85.7%)。
- 适合研究复杂推理任务的AI方法论,尤其关注模型能力释放策略。
国际数学奥林匹克竞赛(IMO)被公认为全球高中生数学最高水平赛事,其题目以难度高、创意性强著称,需深刻洞察力与严谨推理。尽管大语言模型在多数数学基准测试中表现良好,但在奥数级问题上仍常失败。本文通过精心设计的提示工程,构建了一种模型无关的验证与精炼流水线。我们在2025年最新IMO题目上验证该方法的有效性,确保模型未接触比赛数据。使用Gemini 2.5 Pro、Grok-4或GPT-5中的任意一款,经该流程后,正确解答了6道题中的5道(约85.7%准确率)。相比之下,这些模型在32个候选解中选择最优解的基线准确率仅为:Gemini 2.5 Pro为31.6%,Grok-4为21.4%,GPT-5为38.1%。显著提升表明,先进AI推理不仅需要更强的底层模型,更需高效方法论来充分释放其潜力。
原文摘要 · Abstract (English)
The International Mathematical Olympiad (IMO) is widely regarded as the world championship of high-school mathematics. IMO problems are renowned for their difficulty and novelty, demanding deep insight, creativity, and rigor. Although large language models perform well on many mathematical benchmarks, they often struggle with Olympiad-level problems. Using carefully designed prompts, we construct a model-agnostic, verification-and-refinement pipeline. We demonstrate its effectiveness on the recent IMO 2025, avoiding data contamination for models released before the competition. Equipped with any of the three leading models -- Gemini 2.5 Pro, Grok-4, or GPT-5 -- our pipeline correctly solved 5 out of the 6 problems ($\approx$85.7% accuracy). This is in sharp contrast to their baseline accuracies: 31.6% (Gemini 2.5 Pro), 21.4% (Grok-4), and 38.1% (GPT-5), obtained by selecting the best of 32 candidate solutions. The substantial improvement underscores that the path to advanced AI reasoning requires not only developing more powerful base models but also designing effective methodologies to harness their full potential for complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。